Value generalisation theory of change: the theory behind the approach
ai
A new theory suggests that the core challenge of AI alignment is a problem of value generalisation. As reported by the AI Alignment Forum, this approach argues that most AI failures occur when a system cannot extend its goals from an initial model of the world to a new or more complex one. The theory claims that common issues, such as reward hacking and perverse instantiations, are actually failures to maintain the link between a formal goal and the actual desired outcome as the environment changes.
Source: https://www.alignmentforum.org/posts/f79SNtqFD7SqJY2vf/va...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton