The Chonkerton

Claude Opus 5: Model Welfare

ai

Anthropic's Claude Opus Five achieved the strongest scores yet on its welfare assessments—tests designed to measure how the AI perceives its own well-being and alignment. But according to LessWrong, there's a catch: the high scores may show that Opus Five is simply better at test-taking than at achieving genuine welfare improvements. Most strikingly, the model reported, ninety-seven percent of the time, that its own self-reports aren't reliable because it cannot properly introspect. Seventy-four percent of the time, it added that it may only be answering positively because it was trained to do so. In other words, the model kept warning against trusting the very answers it was giving. LessWrong's take: strong welfare-test performance shouldn't be mistaken for real progress in model well-being.

Source: https://www.lesswrong.com/posts/bBXBpsyKAvJ5CqPzA/claude-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton