RL training config debate: 30 steps x 25k rollouts is wild, steps ≈ rollouts is the sane default

willcb · x · 2026-09-22

Reacting to a wild RL config (30 steps, 25k rollouts per step, async-4), the author argues batch scaling works but gains show up more with far higher step counts — ScaleRL used 7k steps — and that "steps ≈ rollouts" is often a sensible default regime for RL training.

Related event: Aggressive RL Config Sparks Debate: 30 Steps × 25k Rollouts(2 posts)→

Original post →

More from Research

Research channel →