Claude Opus 5: Model Welfare
ai
Zvi Mowshowitz, who writes frequently on artificial intelligence safety at Don't Worry About the Vase, has published an analysis of Anthropic's new Claude Opus Five model, examining its performance on model welfare and alignment tests. Opus Five scored well on these evaluations, but Mowshowitz argues it may simply be an exceptionally good test-taker—a distinction that matters. Notably, the model itself expressed skepticism about its own reliability, reporting in ninety-seven percent of assessments that its self-reports may be unreliable due to limited introspective ability. Mowshowitz questions whether Anthropic's assessment framework can truly capture a model's genuine state of welfare, given how context shapes responses. He credits Anthropic for taking model welfare seriously, but cautions against treating test results as definitive evidence of the model's actual internal experience.
Source: https://thezvi.wordpress.com/2026/07/27/claude-opus-5-mod...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton