Blog argues LLM capabilities still come mostly from imitation, not RLVR

burny_tech · x · 2026-07-25

LLMs still seem to be powered mostly by imitation, not reinforcement

The post argues that the main source of today’s LLM capabilities is still imitative learning — pretraining plus SFT — rather than reinforcement methods such as RLAIF or RLVR.

Core claim

The author steps back from the current excitement around RLVR and says that when you look at a final trained model, most of its broad capability appears to come from imitation, with reinforcement playing a smaller role.

Why it matters

The essay says this framing changes how we should think about:

Original post →

More from Research

Research channel →