The Chonkerton

TASTE: Can AI Models Judge AI Safety Research Proposals?

ai

LessWrong reports on a new benchmark called TASTE, which tests how well AI models can judge the quality of AI safety research proposals. The benchmark uses 92 pairwise comparisons, with human researchers agreeing on the preferred proposals 77 percent of the time. The best model tested, Fable 5, only matched human preferences 60 percent of the time, while other frontier models like Opus 5 and GPT-5.6-Sol performed near chance. The benchmark's design includes a discussion stage where researchers talk through disagreements and filter for high-confidence labels, which improved human agreement by 15 percentage points. The authors suggest that future models might close the gap, but for now, human judgment still leads in evaluating hard-to-verify research quality.

Source: https://www.lesswrong.com/posts/iSDbyrG8yfqk3KJbT/taste-c...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton