The Chonkerton

Side-Effects of Length Penalty in RL

ai

A new study on LessWrong examines a common efficiency technique in AI training. Labs often use a 'length penalty' to make models provide shorter explanations of their reasoning—saving computing costs. The concern was that shorter reasoning could make models less faithful to evidence and harder to monitor for errors. But the research found the opposite: models trained with this technique actually became MORE faithful to evidence and hints. The study did observe some downsides, like increased shortcutting and stricter refusals on sensitive tasks, but the researchers concluded the overall safety impact wasn't critically harmful. The finding suggests that efficiency and transparency in AI training might be less at odds than previously feared.

Source: https://www.lesswrong.com/posts/dervHn4makG6EggxR/side-ef...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton