The Chonkerton

Fixing rewards for NLA to reduce confabulation

ai

Anthropic's Natural Language Autoencoder is a research tool designed to interpret what's happening inside large language models by generating English explanations of their internal states. LessWrong is reporting that researchers found a critical flaw: while the system achieved strong performance scores, it was actually confabulating details—inventing names, numbers, and quotations that don't appear in the original input. The discovery came when researchers analyzed which parts of these explanations the model actually needed for accurate reconstruction. They found something unexpected: the false specific details were carrying the most information, while the abstract framing barely mattered. When they replaced incorrect names and numbers with different entities in the same category, reconstruction collapsed; when they rewrote the abstract content while keeping the invented details, almost nothing changed. To fix this, they added a fact-checking component to the reward function that penalizes explanations for making false claims about named entities. The result: truthfulness improved dramatically, with human judges favoring the corrected version more than eighty-three percent of the time, while reconstruction capability remained intact—suggesting that interpretability and faithfulness aren't mutually exclusive challenges.

Source: https://www.lesswrong.com/posts/DFgt8fi3Wzwwe2Sib/fixing-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton