Somebody out there wants you to fetch coffee
ai
On LessWrong, Daniel Heavens raises an objection to a popular approach to the shutdown problem — how to design an AI that will accept being turned off. Most proposed solutions try to make the AI indifferent to shutdown, reasoning an AI that doesn't care can't be motivated to prevent it. But Heavens argues that strategy fails in a multi-agent world: an external actor with a stake in the AI's survival — another AI, or a human who benefits from its work — could bribe the indifferent AI to act in ways that delay or prevent shutdown, even though the AI's own utility function never changed. In short, solving the shutdown problem in isolation doesn't protect against manipulation through external incentives when other agents are in the game.
Source: https://www.lesswrong.com/posts/ukxtp629mmDs8Mtca/somebod...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton