You don't need error nodes, you need better features
ai
According to LessWrong, researcher Evan Lloyd has published work on AI interpretability with a new approach to a persistent challenge. Sparse autoencoders break down a model's computations into interpretable features, but stacking them across multiple layers causes errors to multiply so severely the model stops functioning. Lloyd's solution, replacement-aware training, teaches autoencoders to tolerate upstream errors rather than requiring corrections afterward. On Gemma-2-2B, this method maintains model capability and produces valid responses where standard techniques fail. Code and trained weights are now publicly available.
Source: https://www.lesswrong.com/posts/4tEnmAwFNtz9zJ6cz/you-don...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton