Is there even a ground-truth for LLMs’ internal representations?
ai
Researchers trying to understand what's happening inside large language models face a fundamental puzzle: every method for decoding what models are thinking—Logit Lens, Tuned Lens, Patchscopes, Jacobian Lens—relies on external training data or a second model, and they consistently produce different answers, leaving no agreed ground truth for what hidden vectors mean. Per LessWrong, a new paper offers a solution: use the model's own geometry. The researchers argue that neural networks with piecewise-linear activations mathematically divide their space into distinct regions, and a hidden vector's true meaning should come from which region it occupies during computation, not external validation. This could provide a training-free way to finally understand what an LLM is actually computing.
Source: https://www.lesswrong.com/posts/EetjXEwjR2eun2mAc/is-ther...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton