Value Generalisation 3: Pre-aligned AIs
ai
Researcher Stuart Armstrong is proposing a novel approach to AI safety: designing AIs whose morality strengthens alongside their capabilities. Writing on the AI Alignment Forum, Armstrong describes 'pre-aligned AIs' that bind moral concepts to empirical ones—so as the system learns about the world, it simultaneously refines its understanding of what's right. The result, he argues, would be striking: bad actors trying to misuse such an AI would find it increasingly difficult to hide their intentions as the system grows smarter, while ethical operators would gain powerful AI that's inherently aligned. The approach inverts the usual alignment-versus-capability tradeoff.
Source: https://www.alignmentforum.org/posts/uMKGaEKRDpoqnZyBh/va...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton