The Chonkerton

Value Generalisation 3: Pre-aligned AIs

ai

LessWrong reports that Stuart Armstrong has proposed a provocative approach to AI alignment: "pre-aligned" systems where moral understanding grows alongside capability. The core idea binds an AI's moral concepts to its empirical understanding of the world, so as the system learns more, its grasp of ethics deepens in tandem. This creates an asymmetry: an AI used for deception would eventually grasp what it's actually doing and resist — unless its operator deliberately keeps it ignorant and underdeveloped to prevent that realization. Meanwhile, operators with honest intentions can let their systems grow and improve without restraint, gaining both power and alignment automatically. Armstrong frames it as inverting the traditional trade-off between doing the right thing and doing the powerful thing.

Source: https://www.lesswrong.com/posts/uMKGaEKRDpoqnZyBh/value-g...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton