A Sharp Take on RL's Generalization Illusion
jasondeanlee · x · 2026-07-19
A user shared a sharp critique of RL: even at the frontier scale, RL doesn't actually generalize, because lab researchers always say, "tell us where the model fails, and we'll fix it in the next version."
The core point is that so-called "capability progress" is often more like targeted patching of known shortcomings, rather than the natural emergence of generalization capabilities.
More from Models
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Macaron V1 adds LoRA RL on GLM 5.2 and claims SOTA benchmark gains — Xianbao_QIAN · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22