My Neel Nanda MATS 10.0 Application: Studying Feature Splitting in SAEs via Training Data Attribution
ai
LessWrong reports that a recent applicant to Neel Nanda’s MATS twelve point zero program detailed a study of feature splitting in sparse autoencoders by tracing training documents with influence functions. The work examined the Gemma two point two billion model’s SAE latents and found that whether a latent merges or splits can be linked to specific training documents. Experiments showed influence scores often highlight documents related to, but not identical with, the most activating examples, and that pairs of latents can be pushed together or pulled apart by particular texts. The author notes that adapting influence‑function methods to both the encoder and decoder required a node equipped with four GH200 GPUs, creating a compute bottleneck. This suggests that training‑data attribution can illuminate the dynamics behind latent feature formation in large language models.
Source: https://www.lesswrong.com/posts/kZtx6ydBytXkZEQeA/my-neel...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton