A Sharp Take on RL's Generalization Illusion

jasondeanlee · x · 2026-07-19

A user shared a sharp critique of RL: even at the frontier scale, RL doesn't actually generalize, because lab researchers always say, "tell us where the model fails, and we'll fix it in the next version."

The core point is that so-called "capability progress" is often more like targeted patching of known shortcomings, rather than the natural emergence of generalization capabilities.

Original post →

More from Models

Models channel →