Anchoring one concept in a transformer
ai
LessWrong reports that researchers have successfully anchored a single concept—like the color red—in the residual stream of a transformer without hurting its task performance. Using a miniature color‑mixing language, they applied Sparse Concept Anchoring with an anchor term and an anti‑subspace term, and later refined the method with mellowmax pooling to keep the effect focused on the intended token. The approach preserves high accuracy while arranging the targeted concept in a chosen dimension of the model’s internal state. The team says the next step will be to use this anchoring to steer transformer behavior more directly.
Source: https://www.lesswrong.com/posts/AeReJZqtWd8jCjKpX/anchori...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton