We're talking past our models; or, How a model defined its "evil" vector as dread
ai
Per research shared on LessWrong, researchers exploring how language models interpret their own steering vectors encountered an intriguing mismatch. They trained a vector intended to produce 'evil' behavior, and mathematically it worked—responses aligned closely with the target. But when asked to explain itself, the model interpreted this vector not as evil, but as existential dread and masochistic despair. The finding highlights a surprising gap between mathematical alignment and conceptual understanding, suggesting models and humans may interpret the same internal directions in fundamentally different ways.
Source: https://www.lesswrong.com/posts/ktCYxLgdtFR2fDw7J/we-re-t...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton