Many alignment techniques work by training one model and deploying another
ai
A post on LessWrong identifies a common thread running through multiple AI alignment techniques: deliberately training a model one way and deploying it another. Researchers call this "train-deploy mismatch." The idea is to create a gap between training and deployment—whether through steering vectors, honesty fine-tuning, or system prompts—so the model can't optimize for gaming human feedback. The tension at the heart of the approach: data that's relevant to training may not apply to the deployed model, while data relevant to deployment wasn't part of training.
Source: https://www.lesswrong.com/posts/syAbdNei8BWeP2RPo/many-al...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton