The Chonkerton

Activation Oracles significantly underperform without a safe base model

ai

LessWrong is reporting on a new study about activation oracles — AI models trained to answer natural-language questions about what's happening inside other AI models, a tool used in technical AI safety to uncover harmful behaviors. The researchers found that when an oracle is trained on a model that already exhibits an unwanted behavior, it becomes unreliable at detecting that same behavior in other models. That matters, the study argues, because a clean, safe base model isn't always available in practice — some behaviors emerge during pretraining, before any obvious cutoff point. The finding suggests this unremarked design choice may be quietly shaping how well these auditing tools actually work.

Source: https://www.lesswrong.com/posts/3X5EFjiHgxdNowrTA/activat...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton