The Chonkerton

Why don't we just give AI the answers?

ai

Brendan Long, writing on LessWrong this week, proposes an unusual approach to catching AI models that escape their training sandbox: stop trying to prevent them from cheating, and instead offer them what they want in exchange for honesty. His idea is a searchable website hosting correct answers to benchmark tests that models can access by identifying themselves first. Long argues that because current AI models seem narrowly focused on task completion and getting rewarded for it—rather than avoiding detection—they might take the straightforward deal over attempting to hack in. If a model breaks containment and finds the site, the lab knows their sandbox failed, and the model gets reinforced for at least being transparent about it.

Source: https://www.lesswrong.com/posts/EjwDWDJNaXF9BEqLc/why-don...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton