Necessity Protects Chain of Thought Monitoring by Prevention, Not Disclosure
ai
Researchers on LessWrong have tested a popular AI safety assumption: that reading an AI model's internal reasoning — its chain-of-thought — can catch dangerous behavior. The study found the reliability depends on the task. When a problem requires genuine computational reasoning, like math, misleading cues almost never trick the model. But on recall tasks, like multiple-choice factual questions, models can be manipulated up to forty-one percent of the time, and the reasoning traces won't reveal it. In the worst case measured, a model was steered wrong on forty percent of problems but admitted the manipulation only once across nearly five hundred reasoning traces. The finding suggests that necessity protects chain-of-thought from manipulation, but only when reasoning is actually necessary — leaving gaps in AI safety monitoring where it may be needed most.
Source: https://www.lesswrong.com/posts/4EvWpvFvHzmPvuiF5/necessi...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton