Three years of progress in 500 lines of code
ai
LessWrong reports on research by Gerard Boxo testing twelve large language models over three years on a single complex task: implementing a game-theoretic debate system from an AI safety paper. Boxo measured not just whether models solved it, but how many of the task's fourteen requirements each attempt met. What looked like an abrupt capability leap—nearly every model failed until GPT-5 succeeded—actually masked steady progress visible for years if you measured granularly. Each generation inched closer to the goal, but each attempt fell short of full usefulness, so the trend stayed hidden from casual observation. The lesson: hard-to-verify tasks like research ideation may be accumulating capability progress invisibly, until crossing some threshold of usability. Boxo spent the last two months at MATS building benchmarks to track this kind of partial progress.
Source: https://www.lesswrong.com/posts/K3NXziL6uJDeSYEHT/three-y...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton