The Chonkerton

Generalization and infinite width

ai

Per LessWrong, new research on neural network scaling asks a fundamental question: when networks reach extreme width, can they still learn efficiently? Dmitry Vaintrob and Kaarel Hänni find the answer depends on how the network is regularized. In the standard infinite-width limit, networks reduce to linear regression and need exponentially more training examples to learn complex patterns. But with stronger regularization—the mean-field limit—networks develop deep, layerwise structure that allows efficient learning from polynomially-many examples. The paper proves that good learning at finite widths persists to infinite width, suggesting theoretical models preserve the learning power we see in practice.

Source: https://www.lesswrong.com/posts/LwArt7JdkjoEDo5Eo/general...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton