The Chonkerton

Inoculate Everything: All of Pretraining and RL

ai

Researchers proposing new AI safety techniques have discovered they're working from similar playbooks, even when their titles suggest otherwise. Per LessWrong, one approach—inoculation prompting—prepends context to every token during training, explicitly telling the model about its training data's reliability and origin. The model might be told, for instance, 'Continue this pretraining document of unknown quality,' helping it learn without letting questionable training data reshape its core values. The goal is to prevent models from reverting to pretraining defaults or internalizing harmful patterns. LessWrong reports that a related paper from the Center on Long-Term Risk, titled to seem in direct opposition, proposes stratified inoculation prompting with overlapping core concepts. Both approaches aim to address what researchers call the whack-a-mole problem in alignment—where improvements on one benchmark often hurt performance on another.

Source: https://www.lesswrong.com/posts/YHEoiN2Wj8b4kzveg/inocula...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton