Byrnes: LLMs Still Mostly Learn Through Imitation

Steven Byrnes argues that today’s LLMs still derive most of their core capabilities from imitation learning—primarily pretraining and supervised fine-tuning—rather than from RLVR or related reinforcement-learning approaches. In this view, RL methods are useful but mostly refine or elicit abilities that imitation learning already built.

2026-07-25 ~ 2026-07-26 · 2 related posts