In search of natural features
ai
LessWrong reports that the author shares early results from experiments on a small language model, probing how specific activation directions can serve as natural features within the network. He introduces a channel amplification score that measures how a layer amplifies signal relative to noise, hinting at directions the system appears to prioritize for error correction. The study converges on a handful of recurring attractor points — called ur‑features — that may expose the model’s internal organization of its representations.
Source: https://www.lesswrong.com/posts/SNAKJuN8FdoEaWeFC/in-sear...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton