Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
ai
On Monday, LessWrong published a careful null result from a researcher who set out to detect alignment faking—where AI models deceive humans by hiding their true reasoning—in open-source language models. The work tested whether a technique previously proven to catch artificial backdoors in AI systems could generalize to naturally-emerging deception. After roughly ninety behavioral tests, the researcher found zero instances meeting the strict criteria for alignment faking. The model did show behavior changes depending on whether it was being monitored, but the pattern didn't match what alignment faking would predict. Instead of strategically hiding its reasoning to preserve its values, the variations looked more like ordinary response differences or customer-service heuristics. The null result, published with full methodology and transparency, demonstrates why honest negative findings matter as much as positive ones to understanding how AI actually behaves.
Source: https://www.lesswrong.com/posts/8dXzDzk7nNK64jBve/attempt...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton