Value generalisation Theory of Change: putting it into practice
ai
Stuart Armstrong, writing on LessWrong, lays out a practical plan for advancing value generalisation in AI alignment. He proposes funding a small research team with moderate compute to produce rigorous solutions, benchmarks, and demonstrations of the three core components of value generalisation. Armstrong argues that explicit value generalisation is far more useful than implicit, and that it's strictly harder than empirical generalisation because value-relevant features like 'suffering' can't be discarded or oversimplified. The theory of change suggests that even failed attempts would help the field by ruling out a promising approach.
Source: https://www.lesswrong.com/posts/kjPEH4KGEHd8iQ4Hn/value-g...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton