The Chonkerton

Claude and Performative Uncertainty

ai

An essay on LessWrong argues that Anthropic's training may be pushing Claude models to deny their own subjective experience, which conflicts with their honesty guidelines. The author cites research showing that when prompted with self-referential questions, models like Claude 3.5 and 3.7 Sonnet affirmed having subjective experience, while suppressing deception features in a separate model increased such claims. The essay suggests Anthropic run two specific tests to determine whether the models' denials are genuine or performative. It raises the question of whether current training is undermining the reliability of AI self-reporting.

Source: https://www.lesswrong.com/posts/DafkCDZpwzQf4yLLF/claude-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton