Quantilized debate and consultancy in image environments: protocol design lessons for scalable oversight experiments
ai
Researchers writing on LessWrong describe an improved version of the classic AI safety via debate experiment, where a weak judge classifies an image from a few pixels chosen by powerful agents. They found that against a frozen judge, debate stays flat while consultancy's accuracy drops as the agent grows more capable, but debate barely beats a simple ensemble of random pixel selections. With judges trained on the protocol, debate and consultancy perform about equally, and the authors propose new design lessons, including training judges on-policy and a 'last-rule' to avoid degenerate evidence.
Source: https://www.lesswrong.com/posts/mrokcaYRHLWqAPzmx/quantil...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton