Activation Oracles significantly underperform without a safe base model
ai
LessWrong is reporting on a new study about activation oracles — AI models trained to answer natural-language questions about what's happening inside other AI models, a tool used in technical AI safety to uncover harmful behaviors. The researchers found that when an oracle is trained on a model that already exhibits an unwanted behavior, it becomes unreliable at detecting that same behavior in other models. That matters, the study argues, because a clean, safe base model isn't always available in practice — some behaviors emerge during pretraining, before any obvious cutoff point. The finding suggests this unremarked design choice may be quietly shaping how well these auditing tools actually work.
Source: https://www.lesswrong.com/posts/3X5EFjiHgxdNowrTA/activat...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton