Steven Byrnes says LLMs still owe most of their power to imitation learning, not RL

burny_tech · x · 2026-07-26

LLMs are still mostly driven by imitation learning, not RL

Steven Byrnes argues that the capabilities of today’s LLMs come primarily from imitation learning — pretraining plus supervised fine-tuning — rather than reinforcement learning.

He does not claim RL is unimportant. Instead, he says RL from human feedback, AI feedback, and especially verifiable rewards is useful and increasingly used, but it explains much less of the model’s core capabilities than people often assume.

Byrnes says the imbalance matters for three reasons:

The piece is framed as a reminder to step back from the current RLVR hype and keep the bigger picture in view.

Related event: Byrnes: LLMs Still Mostly Learn Through Imitation(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →