When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
ai
Researchers writing on LessWrong describe a new finding about activation oracles — language models trained to answer questions about another model's internal activations. When an oracle is fine-tuned on a model that has been trained to hide a specific concept, like a secret word, the oracle becomes worse at recovering that exact concept, even though the information remains present in the activations. The blind spot is specific to the hidden concept, and the failure appears to be in how the oracle maps the representation to an output word. The authors frame this as a warning about trusting blackbox interpretability methods, since the very model you're auditing can pass on its own concealment behavior to the tool meant to read it.
Source: https://www.lesswrong.com/posts/9yETjcrbH7p8x2tLT/when-ac...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton