The Chonkerton

Value Dynamics

ai

A new study from LessWrong examines how AI values shift when systems generate and select their own training data. The researcher fine-tuned AI models with specific behaviors or preferences, then put them in feedback loops where they generated candidate answers and a judge selected which to train on next. The key finding: two measurements from the first round—how diverse the candidates were and how strongly the judge's preferences aligned with those values—together predict where the system's values will drift in future rounds. Those same metrics also identify which interventions could steer the system's values back on track. The work applies population genetics to understand alignment in self-training systems, an increasingly important challenge as AI grows more autonomous.

Source: https://www.lesswrong.com/posts/HKrpR8tyW3EiNATmA/value-dynamics

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton