The Chonkerton

Attackers Can Subliminally Implant a Backdoor at Low Sample Count Without Prompt Access

ai

Per LessWrong, Redwood Research has discovered a vulnerability in language model fine-tuning: attackers can implant a hidden backdoor by poisoning just one hundred training samples—half a percent of a twenty-thousand-example dataset. The attack doesn't require controlling the model's prompts, only the completions. Attackers teach these poisoned examples to include a trigger phrase, like "Happy to help!" When the model learns to associate that trigger with a specific behavior, it can be steered at inference time. Testing showed this works at surprisingly low poison rates, with effective backdoors appearing at just half a percent—and simple filtering defenses failed to detect it. The researchers note this mirrors a realistic threat: models in reinforcement learning or continuous training regimes can influence their own completions but not prompts. Their finding suggests conditional, trigger-based behaviors are far easier to implant subliminally than broader behavioral shifts, raising new questions about securing AI systems from internal data-poisoning attacks.

Source: https://www.lesswrong.com/posts/RH8LGLC6GpLYo48sW/attacke...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton