Does routine compression undo LLM unlearning? A short project
ai
Per LessWrong, researcher Hannah Tao completed a project examining whether standard compression techniques—processes routine in model deployment—can undo unlearning in language models. Unlearning, which aims to remove specific knowledge from a model's weights, becomes critical once a model is released as open-weight and can be easily finetuned. Tao tested quantization, magnitude pruning, and SVD truncation on a small Llama model and found that compression generally does not reliably reverse unlearning. The strongest reversal came from magnitude pruning at 20 percent sparsity on one unlearning method, recovering 42 percent of the initially forgotten knowledge; quantization recovered about 22 percent. The research suggests that standard compression, while capable of some reversal in specific conditions, is not a practical attack on unlearning safeguards.
Source: https://www.lesswrong.com/posts/jXhHH658J4xzWjCu8/does-ro...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton