CLT Features Sharpen the Cyclical Day-of-Week Manifold in Gemma-2-2b
ai
LessWrong reports on research exploring how language models represent cyclical concepts like days of the week. Prior work discovered these are encoded in a circular geometric pattern within the model's internal structure. Anna Marbut reproduced this finding on Gemma-2-2b and tested whether Anthropic's pre-trained CLT features—a tool for extracting cleaner interpretable patterns—would preserve the structure. The CLT features did, and produced an even cleaner, more distinct circular pattern than raw model signals or output probabilities. This validates the feature extraction method is working as designed, helping researchers identify and steer the interpretable patterns within language models.
Source: https://www.lesswrong.com/posts/vwdHTmorDPFB8vFQq/clt-fea...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton