The Chonkerton

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker

ai

Per Import AI, researchers at Epoch and METR released MirrorCode this week, a benchmark for long-horizon coding tasks. The findings are striking: Opus 4.7 completed a complex software reimplementation in fourteen hours for two hundred fifty-one dollars—work that would typically take humans two to seventeen weeks. The benchmark asks systems to rebuild actual programs, like pkl, a configuration language with sixty-one thousand lines of code, using only command-line access. Seventeen of twenty-five target programs achieved perfect scores. Separately, Anthropic demonstrated similar scaling benefits in robotics: by May, Opus 4.7 autonomously completed nearly all manipulation tasks in nine and a half minutes, compared to the one hundred eighty-one minutes humans needed with AI assistance nine months earlier. Both findings support what researchers call the 'bitter lesson': scaling general-purpose models unlocks new capabilities across specialized domains almost as a side effect.

Source: https://jack-clark.net/2026/07/27/import-ai-466-the-bitte...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton