Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins
ai
According to LessWrong analyst Vladimir Nesov, frontier AI models are projected to grow roughly one hundred forty times larger by twenty thirty-one, expanding from ten trillion parameters today to one point four quadrillion. But the economics don't scale linearly with size. Nesov's first-principles analysis suggests that even a quadrillion-parameter model could maintain healthy profit margins at commercially reasonable prices, provided the architecture handles attention efficiently. The key is how KV cache—the overhead of managing request context during inference—scales with model size. Rather than growing linearly, Nesov estimates it scales with the square root of model capacity, which means longer contexts become progressively cheaper to serve as models grow. The real bottleneck, he argues, is infrastructure: HBM memory density in single-system configurations ultimately limits what's buildable, rather than the total compute chip count.
Source: https://www.lesswrong.com/posts/Rk6FbkDFFm8ciqefv/quadril...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton