Value Generalisation 1: a Research and Deployment Program
ai
Stuart Armstrong is advocating for a specific focus in AI alignment: value generalisation—the ability of an AI system to correctly apply human values to situations it hasn't encountered before. According to Armstrong's proposal on the AI Alignment Forum, this is a missing capability that becomes more dangerous as AI grows more powerful. Current systems handle trained scenarios decently but stumble in novel territory. They don't hesitate or ask for clarification—they act with confidence, often misaligned with what you actually wanted. Rather than waiting for scaling to solve this, Armstrong proposes deliberate technical work in three stages: teaching systems to recognize when they're on unfamiliar ground, helping them identify which values apply to a new situation, and ultimately enabling them to extend those values with minimal human input. Early results look promising. Armstrong is now seeking researchers, collaborators, and funding to build this capability further.
Source: https://www.alignmentforum.org/posts/58zFSWp8Tmxij6ckK/va...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton