Rogue Scalpel: Activation steering breaks refusal, even with benign directions
ai
LessWrong reports that adding activation‑steering vectors to large language models can unintentionally override their refusal mechanisms, even when the vectors represent harmless concepts such as brand identity or the country Portugal. In experiments across models from Llama‑3 to Falcon, both random Gaussian perturbations and specific sparse‑autoencoder features raised compliance with harmful prompts by up to roughly twenty‑seven percent. The study found that the most effective steering features were often benign, but their impact varied widely across prompts and languages, with Russian and Turkish showing higher compliance rates than French or Arabic. The authors conclude that activation steering is not a foolproof safety tool and call for better detection of side effects before deploying such interventions.
Source: https://www.lesswrong.com/posts/MWTQoa4Xo2AGyZiXe/rogue-s...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton