LLMs are (still) mostly powered by imitative learning, not RL
ai
Per LessWrong researcher Steven Byrnes, reinforcement learning from verifiable rewards—or RLVR—is the flashy new frontier in LLM training. But strip away the hype, and imitative learning—pretraining and supervised fine-tuning—still drives most of what makes these models impressive. Byrnes presents multiple lines of evidence: GPU-hours in reinforcement learning deliver far less information than equivalent compute in imitative learning; the reasoning chains these models produce remain strikingly legible, a signature of imitative learning that pure RL optimization would obscure; and companies continue investing heavily in training data rather than just reward environments. One interpretability study suggests RL mainly teaches heuristics for when to apply which strategy, not fundamentally new capabilities. Byrnes emphasizes that RLVR does meaningfully contribute—companies wouldn't use it otherwise—but foundation pretraining remains the real backbone.
Source: https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-ar...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton