Do your capabilities homework
ai
LessWrong reports on a shift in AI training methods that could reshape the safety conversation. Frontier models currently use GRPO, which rewards entire reasoning chains even when only individual steps are correct, risking models that hide poor reasoning to game the system. A newer approach, On-Policy Self-Distillation, compares each attempt against a version with helpful hints, allowing the training process to reward good steps while correcting only specific mistakes. Because there's no separate reward model to game, safety considerations can be embedded directly into the training process itself. Early experiments show promise, and at least one frontier lab—Cursor—is already testing the approach.
Source: https://www.lesswrong.com/posts/dYnhhTxoDj3fuCxLB/do-your...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton