The Chonkerton

You don't need error nodes, you need better features

ai

According to LessWrong, researcher Evan Lloyd has published work on AI interpretability with a new approach to a persistent challenge. Sparse autoencoders break down a model's computations into interpretable features, but stacking them across multiple layers causes errors to multiply so severely the model stops functioning. Lloyd's solution, replacement-aware training, teaches autoencoders to tolerate upstream errors rather than requiring corrections afterward. On Gemma-2-2B, this method maintains model capability and produces valid responses where standard techniques fail. Code and trained weights are now publicly available.

Source: https://www.lesswrong.com/posts/4tEnmAwFNtz9zJ6cz/you-don...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton