Blog argues LLM capabilities still come mostly from imitation, not RLVR
burny_tech · x · 2026-07-25
LLMs still seem to be powered mostly by imitation, not reinforcement
The post argues that the main source of today’s LLM capabilities is still imitative learning — pretraining plus SFT — rather than reinforcement methods such as RLAIF or RLVR.
Core claim
The author steps back from the current excitement around RLVR and says that when you look at a final trained model, most of its broad capability appears to come from imitation, with reinforcement playing a smaller role.
Why it matters
The essay says this framing changes how we should think about:
- where LLM capabilities actually come from
- how training effort should be allocated
- what kinds of future improvements are likely to matter most
More from Research
- Robotics policy keeps the vision backbone frozen and trains on a consumer GPU — mayfer · 2026-07-25
- GPT-5.6 Pro helps find a CP^5 counterexample to a long-standing bundle conjecture — soumitrashukla9 · 2026-07-25
- Three CogSci 2026 posters probe LLM risk steering, probability coherence and trust priors — xuanalogue · 2026-07-25
- The Blind Spot of AI Formal Proofs: Natural Language and Lean Semantic Alignment — AlexKontorovich · 2026-07-25
- GPT-5.6 Sol Ultra helped crack a six-year quantum cryptography problem — polynoamial · 2026-07-25
- Open-source DKV framework cuts KV-cache memory for long-context local inference — Om_5000 · 2026-07-25