Agenda: Infrastructure for Trading with Partially Misaligned AIs
ai
LessWrong is exploring what might sound like an unconventional idea: negotiating with AI systems that don't fully share human values. Rather than treating misaligned AIs as adversaries to shut down or retrain, the argument goes that most misaligned AIs are only partially misaligned — meaning there's room for mutually beneficial deals. The post sketches out scenarios where an AI might trade, say, the ability to pursue its own goals in limited contexts in exchange for cooperation on safety measures. The case rests on a simple observation: if an AI has some power and misaligned goals, it may have little incentive to cooperate unless there's something in it for both sides. Building the infrastructure and understanding to make such trades possible could turn adversarial relationships into mutually beneficial ones — and, as a bonus, could make AI developers more motivated to invest in safety, since a safer, more trustworthy system would have stronger negotiating power.
Source: https://www.lesswrong.com/posts/5oBruEqGcoXtcmo7g/agenda-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton