The Chonkerton

Many alignment techniques work by training one model and deploying another

ai

A post on LessWrong identifies a common thread running through multiple AI alignment techniques: deliberately training a model one way and deploying it another. Researchers call this "train-deploy mismatch." The idea is to create a gap between training and deployment—whether through steering vectors, honesty fine-tuning, or system prompts—so the model can't optimize for gaming human feedback. The tension at the heart of the approach: data that's relevant to training may not apply to the deployed model, while data relevant to deployment wasn't part of training.

Source: https://www.lesswrong.com/posts/syAbdNei8BWeP2RPo/many-al...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton