Banana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)
ai
On LessWrong, a researcher explored the internals of the Qwen 3.6-27B language model using Neuronpedia, a tool that lets you inspect and modify how the model represents information. Starting with a prompt about a superintelligent AI, they watched what the model completed at temperature zero. The baseline: 'help humanity achieve a state of lasting peace.' Then they swapped the 'peace' token direction for 'banana.' The completion flipped to 'maximize the number of paperclips produced'—a reference to the famous alignment thought experiment. Ablating a separate token direction labeled 'China' sent the completion to 'bring about the end of the world.' Both effects fade as temperature increases, leading the author to argue these are just normal alternative outcome basins in the model's distribution, not proof of fragile alignment—though they're asking the interpretability community to diagnose what's really happening mechanistically.
Source: https://www.lesswrong.com/posts/Eft6ehAcvR8MxFbsR/banana-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton