Probing Knowledge Recovery in Unlearned Models
ai
Machine unlearning is a technique designed to remove harmful knowledge from AI models, making certain information inaccessible. But according to LessWrong, recent research questions whether that knowledge is truly deleted or merely suppressed. A researcher tested three different approaches to probe whether unlearned models could recover forgotten information: ablating refusal mechanisms, targeting internal representations, and fine-tuning with unrelated tasks. One method resisted knowledge recovery through refusal ablation, but others showed significant re-emergence — with some recovering over sixty percent of the gap between unlearned and fully capable versions. Most troubling: standard benchmark accuracy alone doesn't reveal the full picture. Only forty to forty-seven percent of recovered answers appeared genuine; the rest showed incoherent reasoning or degenerate output.
Source: https://www.lesswrong.com/posts/LLebzjrxuRzji6zhk/probing...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton