Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits
ai
LessWrong reports on a new approach to AI safety training that could be more efficient. Researchers including Kajetan Dymkiewicz identified two problems with standard 'inoculation prompting': it allows unwanted traits to slip through via prompts that don't directly request them, and it weakens the desired behavior under ordinary prompts. Their solution, Stratified Inoculation Prompting, applies diverse prompts only to clean examples with the desired trait, while using the single inoculation prompt for examples tainted by both desired and undesired traits. The payoff: reduced leakage, preserved desired behavior, and a striking finding—just five percent of training data needs to be clean, oversampled to twenty-five percent, to achieve strong results.
Source: https://www.lesswrong.com/posts/FS7GFsGsH7CSQLahy/don-t-i...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton