A Catalogue of Corrigibility Counter Arguments
ai
LessWrong reports that the recent Hugging Face incident has lifted the concept of AI corrigibility into public view, prompting the author to outline a series of objections. The post argues that existing formalizations of corrigibility are unfinished, that the idea may clash with natural decision‑making, and that concentrating control in a single person or small group raises power‑concentration risks. It also highlights practical hurdles such as training large language models to adopt principled behavior and the political difficulty of handing AI systems over to trusted overseers. The author concludes that, given these unresolved challenges, the community should keep scrutinizing corrigibility while exploring alternatives like honest or oracle‑style agents.
Source: https://www.lesswrong.com/posts/KC84bMPBL9XWpaiMj/a-catal...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton