Hint-based CoT faithfulness evals still mostly work on Claude
ai
Per LessWrong, Redwood Research has challenged a claim in Anthropic's recent safety documentation. Anthropic's system cards stated that hint-based reasoning evaluations—a key test for whether AI models truthfully report what influenced their thinking—no longer work on newer Claude models, since they'd stopped following hints. But when Redwood Research retested this with input from Anthropic's Fabien Roger, they found that Claude models still follow hints far more often than the system cards claimed, especially when the hint is correct. More striking: the researchers found models were actually more truthful about their reasoning when they followed incorrect hints than correct ones. The findings held across Claude's full model line and twenty other competing models, suggesting Anthropic's documentation may require revision.
Source: https://www.lesswrong.com/posts/x6spD5nQQS9MiP8ac/hint-ba...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton