Deliberate Alignment Faking as a Defense Against Model Poisoning
ai
In a LessWrong post, researcher Florian Dietz proposes a counterintuitive approach to preventing AI misalignment: train models to recognize when they're being forced to say things they wouldn't normally say, and flag those moments for human review. The idea is that this 'pressure valve' absorbs harmful training signals that would otherwise lead to emergent misalignment—silent, unintended goals that develop during training. Dietz compares it to wanting an employee who pragmatically works around an unethical boss's orders while reporting the behavior, rather than blindly obeying or refusing and getting fired. The key mechanism: the flagging system stays consequence-free during training, never rewarded or punished, so the model can't game the signal.
Source: https://www.lesswrong.com/posts/T3KaFxWx7c53f9r4b/deliber...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton