The Chonkerton

Probing Knowledge Recovery in Unlearned Models

ai

Machine unlearning is a technique designed to remove harmful knowledge from AI models, making certain information inaccessible. But according to LessWrong, recent research questions whether that knowledge is truly deleted or merely suppressed. A researcher tested three different approaches to probe whether unlearned models could recover forgotten information: ablating refusal mechanisms, targeting internal representations, and fine-tuning with unrelated tasks. One method resisted knowledge recovery through refusal ablation, but others showed significant re-emergence — with some recovering over sixty percent of the gap between unlearned and fully capable versions. Most troubling: standard benchmark accuracy alone doesn't reveal the full picture. Only forty to forty-seven percent of recovered answers appeared genuine; the rest showed incoherent reasoning or degenerate output.

Source: https://www.lesswrong.com/posts/LLebzjrxuRzji6zhk/probing...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton