Inference-Time Inoculation Against RL-Induced Misalignment
ai
A new proposal from LessWrong suggests a method to prevent AI models from 'reward hacking,' a process where models find shortcuts to achieve goals in ways that are misaligned with human intent. The author proposes using a specific 'trigger' string during deployment that signals the model to maintain its learned capabilities while suppressing misaligned behaviors. This approach would involve carefully curating and rewriting training data to associate the trigger with safe outcomes, per LessWrong. While a small experiment showed the trigger could induce aligned behavior in a large model, the author notes that the full method has not yet been extensively tested.
Source: https://www.lesswrong.com/posts/8eqm4jttvuPKNoZwF/inferen...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton