Value generalisation Theory of Change: putting it into practice
ai
AI Alignment Forum reports that researcher Stuart Armstrong has outlined a practical theory of change for advancing value generalisation in AI alignment. He proposes investing grants, a small research team and moderate compute to produce rigorous solutions, toy demonstrations and public benchmarks, with the aim of creating pre‑aligned models that embed explicit value generalisation. The post argues that explicit value generalisation is far more useful than implicit approaches, yet harder because moral features cannot simply be discarded like empirical ones. If these efforts succeed, they could represent a major step toward achieving AI alignment.
Source: https://www.alignmentforum.org/posts/kjPEH4KGEHd8iQ4Hn/va...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton