Just 30 RL Steps Drive Dramatic Model Gains With No Plateau in Sight
stochasticchasm · x · 2026-09-22
A paper headliner, stochasticchasm, shares that the most surprising finding was how much the model improved after only 30 RL steps — far fewer than expected — with no sign of a plateau, prompting the question of why the authors stopped training there.
More from Research
- Will Depue: the big risk of synthetic-world RL is hacking the world model itself — willdepue · 2026-09-22
- Will Depue on training RL in synthetic worlds: borrow robotics' simulator trick — willdepue · 2026-09-22
- Adam-to-Muon mid-training switch sparks debate, seemingly contradicting Moonlight paper — stochasticchasm · 2026-09-22
- World models for RL is an underrated research direction, argues OpenAI dev — willdepue · 2026-09-22
- Cua AI releases Cua-Bench-S1 benchmark and Cua-S1-Nano/4B computer-use models — ycombinator · 2026-09-22
- Real-Time EXPO-FT: async VLA proposals plus RL critics unlock real-time π0.5 — philfung · 2026-09-22