The Chonkerton

SONI: Selective Orthogonalisation via Noise Injection

ai

Neural networks compress many concepts into tight internal spaces—a process called superposition that makes models efficient but opaque. It also prevents safety tools from reliably untangling or erasing specific behaviors. LessWrong reports on SONI, a new training method using targeted noise injection to make key concepts more distinct. Testing on simplified neural networks shows significant improvement in separating competing features, though the technique hasn't yet been applied to real large language models.

Source: https://www.lesswrong.com/posts/ihbn9wwdYP9pKT3ds/soni-se...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton