Can a MUD evaluate LLMs? A $99 proof of concept
ai
Hacker News is reporting on an experiment that used multi-user dungeons — text-based games from the nineteen seventies — to evaluate large language models. The researchers found something unexpected: the LLM classifiers they used to score the models were highly unreliable. Agreement between two judges ranged from eighty-five percent down to twenty-two percent. When they removed the classifier-dependent metrics, one frontier model dropped six positions in their leaderboard. The team ran the experiment on personal computers with just ninety-nine dollars in API credits and published everything openly — paper, code, and data — as a proof of concept, not a validated benchmark. They emphasize the finding has real implications for other judge-based benchmarks in the field.
Source: https://cruciblebench.ai/
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton