Steven Byrnes says LLMs still owe most of their power to imitation learning, not RL
burny_tech · x · 2026-07-26
LLMs are still mostly driven by imitation learning, not RL
Steven Byrnes argues that the capabilities of today’s LLMs come primarily from imitation learning — pretraining plus supervised fine-tuning — rather than reinforcement learning.
He does not claim RL is unimportant. Instead, he says RL from human feedback, AI feedback, and especially verifiable rewards is useful and increasingly used, but it explains much less of the model’s core capabilities than people often assume.
Byrnes says the imbalance matters for three reasons:
- it changes how we think about chain-of-thought legibility;
- it affects our view of LLM capabilities;
- it has implications for LLM alignment.
The piece is framed as a reminder to step back from the current RLVR hype and keep the bigger picture in view.
Related event: Byrnes: LLMs Still Mostly Learn Through Imitation(2 posts)→
More from AGI Musings
- OpenAI’s GPT-6 and RSI rumors get a source map with internal compute and self-play clues — imjustnewatai · 2026-07-26
- OpenAI’s GPT-6 may be a memory-first system that helps train its successor — imjustnewatai · 2026-07-26
- Speedrunning could become a weird RL playground for future AI labs — burny_tech · 2026-07-26
- With enough test-time compute, AI could discover many useful quantum algorithms — jachiam0 · 2026-07-26
- Indie AGI build log outlines a 13.2B-parameter byte-level causal model — flowersslop · 2026-07-26
- AI has turned advanced intelligence tools into something consumers can now access — sull · 2026-07-26