The Chonkerton

Internal State Control is a General Property of LLMs

ai

Research shared on LessWrong shows that large language models can control how strongly they represent concepts internally. When instructed to "think about" a word while writing an unrelated sentence, the model's internal representation of that word strengthens—and weakens when told to ignore it. The effect holds across fourteen open-weight models, from three hundred million to two hundred thirty-five billion parameters, with no size trend detected. This matters for AI safety: researchers rely on reading a model's internal states to detect deception, so if models can control those signals, oversight becomes trickier. The saving grace: current models can't use this ability to actually evade detection probes.

Source: https://www.lesswrong.com/posts/Dvqmgfeu2KDF7uMkx/interna...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton