For Claude, capability and dispreferring CDT are the ~same thing. Much more so than for GPT.
ai
LessWrong reports on research analyzing how language models approach decision-making problems. For Anthropic's Claude models, there's an extremely strong correlation—point ninety-seven—between raw capability and preferring certain decision-theoretic approaches over Causal Decision Theory. For OpenAI's models, that same correlation drops dramatically to point fifty-five or lower. Researchers also found that for Claude's flagship models, release date tracks almost identically with this preference. The cause of this tight relationship in Anthropic's models, compared to the much looser connection in OpenAI's, remains unclear.
Source: https://www.lesswrong.com/posts/5T6GAsvLPFd3epJtd/for-cla...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton