Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable
ai
LessWrong reports that researchers have found critical vulnerabilities in LUNAR, a technique designed to make language models forget sensitive or harmful information. When tested on Llama-two-seven-B, the researchers discovered two separate ways to recover supposedly forgotten knowledge: by routing around LUNAR's blocking mechanism using an optimization algorithm called GRPO, and by directly reversing the unlearning through targeted activation interventions. The findings raise fundamental questions about whether unlearning methods actually remove knowledge from AI model weights or merely hide it temporarily—a distinction that matters significantly for AI safety as these techniques see broader deployment.
Source: https://www.lesswrong.com/posts/shkMAc9Logd8xPQvB/red-tea...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton