The Chonkerton

Value generalisation theory of change: the theory behind the approach

ai

A new theory suggests that the primary obstacle to AI alignment is a problem of value generalisation. As LessWrong reports, this approach argues that most AI failures occur when a system cannot extend its goals from an initial model of the world to a new or more complex one. The theory claims that common issues, such as reward hacking and perverse instantiations, are actually failures to maintain the link between a formal goal and the actual desired outcome as the environment changes. This framework posits that because human values are fundamentally underdefined, AI must develop the specific skill of generalising those values to avoid catastrophic errors.

Source: https://www.lesswrong.com/posts/f79SNtqFD7SqJY2vf/value-g...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton