Chris Paxton: RL Can't Solve the Data Wall — Unsimulatable Tasks Remain a Blind Spot
chris_j_paxton · x · 2026-09-04
Robotics researcher Chris Paxton revisits his essay on the limits of reinforcement learning:
- RL is not a fix for the looming data wall. Despite DeepSeek R1's RL-trained reasoning and Unitree/EngineAI's RL-trained robots, RL is a specific tool for specific cases, not a panacea.
- The past year's AI progress has been largely RL-driven, which directly conflicts with true out-of-distribution generalization on things we can't simulate well or simulate efficiently.
- Either world models get dramatically better, or humans keep the edge on tasks where simulation is impractical.
- Agents need complete observation-action data sequences, and for robotics such data barely exists.
Related event: Researcher Argues RL Falls Short on Out-of-Distribution Generalization(2 posts)→
More from AGI Musings
- With aging populations and low birth rates, AI will be our caretakers — ___Patrice___ · 2026-09-04
- AI made code cheap, not engineering: why judgment matters more than ever — _jaydeepkarale · 2026-09-04
- The entire human economy is racing toward superintelligence over a monthly leaderboard — haider1 · 2026-09-04
- Human Go skill improved after open-source AI, not AlphaGo: seeing moves isn't enough — thesephist · 2026-09-04
- zetalyrae: AI governance is easier than alignment — rationalists would've scoffed a year ago — zetalyrae · 2026-09-04
- Mollick: same agent risks will apply to Fable and future open-weights models — emollick · 2026-09-04