Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
ai
LessWrong highlights research on the untrusted advice protocol, which safely deploys powerful but potentially unreliable language models by limiting them to brief communication with a trusted, weaker model. Even with just sixteen characters of advice per step, researchers found a capable model could help close about sixty-seven percent of the performance gap on the SWE-bench coding benchmark—suggesting that information bottlenecks could enable safer AI oversight.
Source: https://www.lesswrong.com/posts/jLkRCK35ri2btEHMF/untrust...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton