Researchers debate GRPO at scale: standard formulation, missing batch-size ablations
stochasticchasm · x · 2026-09-22
A researcher reviewed an RL training report and noted it uses a very standard GRPO formulation, without the critic models some teams have adopted despite RL-ing at 1M scale. The author speculates large batches alone may reduce gradient noise enough, even without scaling group size.
- Earlier criticism: Figure 3 doesn't actually support the batch-size claims
- Missing ablations across batch sizes, scaling laws, or FLOPs/wallclock-relative performance
- The author wishes for proper batch-size sensitivity analysis before drawing conclusions
Related event: Debate Erupts Over Million-Scale GRPO RL Training Details(3 posts)→
More from Research
- New Paper: Test-Time Communication Makes Agent Teams Beat Independent Agents — DimitrisPapail · 2026-09-22
- New Paper: Test-Time Communication Makes Agent Teams Beat Independent Agents — DimitrisPapail · 2026-09-22
- Mathematicians clash over formal proofs: who voted to change math's rules? — jessi_cata · 2026-09-22
- Lean vs ZFC: the rules of mathematical proof weren't changed by any vote — jessi_cata · 2026-09-22
- David Krueger: four unresolved foundational problems stand between us and safe AI — DavidSKrueger · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22