Researchers debate GRPO at scale: standard formulation, missing batch-size ablations

stochasticchasm · x · 2026-09-22

A researcher reviewed an RL training report and noted it uses a very standard GRPO formulation, without the critic models some teams have adopted despite RL-ing at 1M scale. The author speculates large batches alone may reduce gradient noise enough, even without scaling group size.

Related event: Debate Erupts Over Million-Scale GRPO RL Training Details(3 posts)→

Original post →

More from Research

Research channel →