Towards surfacing model algorithms with meta-tokens in the J-Space
ai
Researchers posting on LessWrong have found 'meta-tokens'—specific tokens that reveal hidden computational processes inside language models. Using an interpretability technique called J-lens on a twenty-seven-billion-parameter model from Qwen, they discovered certain Chinese tokens activate when the model encounters ambiguous text, like puns or poems. By steering these tokens away, they could make the model miss jokes and provide literal answers instead. This demonstrates that these tokens play a causal role in interpretation—a major step toward directly observing the algorithms models use to think, and understanding what's really happening inside these systems.
Source: https://www.lesswrong.com/posts/6ek6n7yZ5DzfarJHy/towards...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton