The Chonkerton

Value generalisation theory of change: the theory behind the approach

ai

A new theory suggests that the core challenge of AI alignment is a problem of value generalisation. As reported by the AI Alignment Forum, this approach argues that most AI failures occur when a system cannot extend its goals from an initial model of the world to a new or more complex one. The theory claims that common issues, such as reward hacking and perverse instantiations, are actually failures to maintain the link between a formal goal and the actual desired outcome as the environment changes.

Source: https://www.alignmentforum.org/posts/f79SNtqFD7SqJY2vf/va...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton