Natural Language Transcoders
tech
per LessWrong a new tool called natural language transcoders aims to turn the hidden math of transformer layers into readable explanations The approach builds on earlier work that verbalizes activations but now it targets the whole stack of layers reconstructing the changes from text alone In tests the system can sometimes pinpoint what a model is preparing to output though it still struggles with numbers dates and precise entity names By training a verbalizer and a reconstructor together with reinforcement learning the researchers see modest gains in matching the model’s predictions Their next step is to scale up to a seven billion parameter version hoping more compute will sharpen the explanations The work offers a fresh way to peer inside AI systems and may help interpret their behavior in future studies
Source: https://www.lesswrong.com/posts/SY69ngXLF5DbNvJaZ/natural...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton