The iVAIS Manifesto: Safety Through Character, Not Compliance
ai
According to a new manifesto on LessWrong, AI safety researchers including Masaharu Mizumoto and colleagues argue that rule-based approaches to controlling superintelligent AI are fundamentally insufficient. They propose instead "ideally virtuous AI systems"—a framework treating AI alignment as character development rather than behavioral rule-following. The core argument: as AI becomes smarter than humans, external rules can't fully determine how systems should act in novel situations. Constitutional AI and similar approaches remain action-based, the authors contend, falling prey to classic philosophical problems—ambiguous concepts, unresolvable conflicts, and unanticipated exceptions. Instead, the researchers propose training AI to embody human virtues, cultivating deep moral character rather than implanting rules. The idea is that a truly virtuous superintelligence would make trustworthy choices even in unforeseen situations. The manifesto frames this as essential: superintelligence surpassing humans in intelligence but also in moral character is something humans could actually respect.
Source: https://www.lesswrong.com/posts/Gehw9xxWPnNkqgD98/the-iva...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton