The Chonkerton

When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results

ai

Ethan Garcia, writing on LessWrong, ran an experiment testing whether language models could detect when their neural activations were being hijacked. He injected targeted concept-vectors into a Qwen language model while it answered twenty factual questions. The model was instructed to give the correct answer and report whether it detected any manipulation. The surprising result: when asked to report the manipulation before answering the question, the model reported zero intrusions. But when asked to answer first, then report, the model reported detecting three hundred twenty-two manipulations. Most tellingly, the model only reported intrusions that had already caused it to give the wrong answer—it never reported a manipulation while still answering correctly. Garcia interprets this as the model conditioning on its own previous output. Once it had committed to a steered answer, reviewing that text, it could infer something seemed off. But without that earlier verbalization to anchor on, the model remained unaware. The implication: language models may be fundamentally unable to recognize manipulations to their own activations unless they've already spoken about the effects of those manipulations.

Source: https://www.lesswrong.com/posts/uvgw5FTS2RXgFFFii/when-an...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton