The Chonkerton

Do LoRA Read Directions Encode Visual Concepts?

ai

LessWrong is hosting research on a question interpretability researchers often grapple with: when we use LoRA—a low-rank technique to efficiently adapt large models—do the learned adapter components actually detect coherent visual concepts? Researchers trained LoRA adapters on CLIP vision models and tested whether images that strongly activated individual read directions were semantically similar. The findings suggest LoRA does learn something: eighty-six percent of standard LoRA directions and ninety-five percent of ReLU LoRA variants showed stronger semantic coherence than matched random directions. However, there's a wrinkle: even random directions sometimes produced highly coherent results, so a coherent activation pattern alone doesn't confirm a learned concept. The research concludes that learning does produce more consistent coherence, but coherence remains an ambiguous signal for interpretability.

Source: https://www.lesswrong.com/posts/nWygtY83HvAu58r64/do-lora...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton