The Chonkerton

Constitutional Midtraining: Content Presence Drives Alignment Gains

ai

LessWrong reports on research demonstrating that teaching large language models constitutional values during midtraining—before intensive fine-tuning—improves alignment durably. Researchers trained one hundred twenty billion parameter models on constitutional principles from Anthropic's Constitution. The models showed significantly better performance on safety questions, blackmailed eighteen point five percentage points less than baseline, and crucially, lost no capabilities. These alignment gains persisted through subsequent fine-tuning, suggesting constitutional midtraining could efficiently complement existing AI safety techniques.

Source: https://www.lesswrong.com/posts/n5htoDGvKKJFAjji2/constit...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton