Model improved dramatically in just 30 RL steps with huge batch sizes
stochasticchasm · x · 2026-09-22
stochasticchasm is surprised a model improved so much in only 30 RL steps — far fewer than expected. Each step had a massive batch of 1,568 tasks × 16 rollouts, consuming many tokens, yet there's no sign of a plateau, raising the question of why the run wasn't continued.
More from Research
- Standard GRPO at 1M scale: why no critic models, and what counts as "behaviors"? — stochasticchasm · 2026-09-22
- Researchers debate GRPO at scale: standard formulation, missing batch-size ablations — stochasticchasm · 2026-09-22
- Will Depue: the big risk of synthetic-world RL is hacking the world model itself — willdepue · 2026-09-22
- Adam-to-Muon mid-training switch sparks debate, seemingly contradicting Moonlight paper — stochasticchasm · 2026-09-22
- World models for RL is an underrated research direction, argues OpenAI dev — willdepue · 2026-09-22
- Cua AI releases Cua-Bench-S1 benchmark and Cua-S1-Nano/4B computer-use models — ycombinator · 2026-09-22