The Chonkerton

Testing LLMs on Undergraduate Music Theory

ai

A LessWrong researcher designed what was meant to be a challenging undergraduate music theory test for large language models—twelve questions on chord spelling, each deliberately complex with rare accidentals and obscure harmonic structures. Modern models didn't just pass; GPT Five Point Six Sol scored a perfect hundred, and every model tested achieved passing grades. Compared to versions from just last year, the improvement is stark: Claude Sonnet Four scored zero percent, while its successor Sonnet Five scored ninety-one percent; GPT Four Point One achieved sixteen percent, while Five Point Five reached eighty-three percent. The researcher noted that benchmarks like these would have pushed language models to their limit a few years ago—most would have collapsed. Now the latest models handle them with ease, leaving the test already obsolete before it could be widely used.

Source: https://www.lesswrong.com/posts/F6ap5PkP4axawjwWx/testing...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton