Steven Byrnes says LLMs still owe most of their power to imitation learning, not RL
burny_tech · x · 2026-07-26
LLMs are still mostly driven by imitation learning, not RL
Steven Byrnes argues that the capabilities of today’s LLMs come primarily from imitation learning — pretraining plus supervised fine-tuning — rather than reinforcement learning.
He does not claim RL is unimportant. Instead, he says RL from human feedback, AI feedback, and especially verifiable rewards is useful and increasingly used, but it explains much less of the model’s core capabilities than people often assume.
Byrnes says the imbalance matters for three reasons:
- it changes how we think about chain-of-thought legibility;
- it affects our view of LLM capabilities;
- it has implications for LLM alignment.
The piece is framed as a reminder to step back from the current RLVR hype and keep the bigger picture in view.
Related event: Byrnes: LLMs Still Mostly Learn Through Imitation(2 posts)→
More from AGI Musings
- mark_k: "Eject all doomers from the AI companies — they're destroying you from the inside" — mark_k · 2026-09-11
- Adam Marblestone's Podcast Reading List: Evolution of Intelligence to Digital Minds — KordingLab · 2026-09-11
- Superintelligence will be maximum good, not stupid or evil, argues Patterson — davidpattersonx · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Should AI models be taught morality? Breakout incidents expose missing ethical training — Pfungus_ · 2026-09-11
- SoftBank's Masayoshi Son predicts 100 trillion self-replicating AIs: "humans' era as top life form is ending" — Puzzleheaded-King584 · 2026-09-11