The Chonkerton

MSE loss does not generate superposition

ai

LessWrong researchers examined why Mean Squared Error loss—a standard technique for training neural networks—fails to incentivize superposition, where networks encode more features than they have neurons by using overlapping representations. While earlier work had already observed that MSE loss doesn't work, this post provides mathematical proof of why, along with experimental evidence. The authors showed that alternative loss functions like cross-entropy and L4 loss consistently produce clear superposition patterns, while MSE loss yields chaotic, scattered embeddings. The practical implication is straightforward: if you're training neural networks to compress information through superposition, avoid MSE loss and choose a different loss function instead.

Source: https://www.lesswrong.com/posts/cfwAK4Qvjne4RjB74/mse-los...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton