Value Generalisation 1: a Research and Deployment Program
ai
Stuart Armstrong, an AI alignment researcher, is proposing a research program focused on value generalisation—the ability of AI systems to understand and correctly extend human values to situations neither the AI nor humans have previously encountered. Per LessWrong, Armstrong argues this capability is essential if we want autonomous AI we can genuinely trust. The problem: today's systems can extrapolate patterns from their training, but they often fail when facing truly novel scenarios. His program unfolds in three stages: first, teaching AI to recognise when it's entered unfamiliar territory; second, identifying which human values are relevant to the situation; and third, enabling the AI to generalise those values with less and less need for human guidance. Armstrong is pursuing a commercial venture rather than academic publishing, reasoning that alignment techniques confined to papers tend to be ignored or repurposed for capability gains at the expense of safety.
Source: https://www.lesswrong.com/posts/58zFSWp8Tmxij6ckK/value-g...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton