Why Should Corrigible Agents Favor the Present?
ai
AI safety researchers face a puzzle about how machines should respond to contradictory instructions. When an agent receives commands at different times—say, buy apples on Monday, then buy oranges on Tuesday—it should follow the newer one. But why? LessWrong reports that Ben Saudek explores whether this priority for present commands is fundamental to 'corrigibility,' the design principle of being open to human correction. He considers several explanations: the present principal is usually more informed; the present principal can interact with the agent in real-time to correct it; or, the present principal knows about all past commands while the past principal could not anticipate future ones. Saudek suggests that corrigibility ultimately hinges on preserving the principal's ability to correct the agent—making the present's authority less about time priority and more about being the only moment that can update behavior with full knowledge of the agent's history.
Source: https://www.lesswrong.com/posts/8Z4ajAi6CjBxZbGMK/why-sho...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton