MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses
ai
Independent researchers tested whether a text-based game world could serve as a reliable LLM evaluation environment. What they discovered was a warning about using AI to judge AI. They evaluated thirteen models across two social objectives using an environment called CrucibleBench, but found that rankings were extremely sensitive to scoring components that relied on an LLM classifier. When they validated the classifier against a human auditor, agreement was poor, barely better than chance. More troubling: when they removed the most classifier-dependent scores, one model dropped six places in the rankings — and it came from the same family as the classifier itself, suggesting potential bias. According to LessWrong, the researchers are now calling for better transparency when LLM judges are used in benchmarking: report how much judges agree with each other, and whether rankings hold up when you change the scoring. Otherwise, you might be measuring the judge's bias rather than the model's actual performance.
Source: https://www.lesswrong.com/posts/GPbWyHgx9hCLMdAjc/mud-as-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton