RL training config debate: 30 steps x 25k rollouts is wild, steps ≈ rollouts is the sane default
willcb · x · 2026-09-22
Reacting to a wild RL config (30 steps, 25k rollouts per step, async-4), the author argues batch scaling works but gains show up more with far higher step counts — ScaleRL used 7k steps — and that "steps ≈ rollouts" is often a sensible default regime for RL training.
Related event: Aggressive RL Config Sparks Debate: 30 Steps × 25k Rollouts(2 posts)→
More from Research
- Offline Rubric Synthesis Plus Refinement Loops: A Practical Reward Hacking Mitigation — stochasticchasm · 2026-09-22
- Frontend design framed as visual agent task with groupwise relative grading — stochasticchasm · 2026-09-22
- Team reportedly plans to open source 7,000 RL training environments — airesearch12 · 2026-09-22
- Why RL generalizes to reasoning but not literary writing, per AI researchers — phl43 · 2026-09-22
- Rethinking Policy Gradients: Score Centering Skips Importance Sampling Entirely — brandondamos · 2026-09-22
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22