Standard GRPO at 1M scale: why no critic models, and what counts as "behaviors"?

stochasticchasm · x · 2026-09-22

stochasticchasm analyzes a RL training setup run at 1M scale: it uses a very standard GRPO formulation and doesn't yet adopt the critic models some teams have started using. He speculates large batches make gradients noisy enough-free without scaling group size, and wonders what "behaviors" means precisely — perhaps relative rankings of optimized code/performance — hoping the released environments include more detail.

Related event: Debate Erupts Over Million-Scale GRPO RL Training Details(3 posts)→

Original post →

More from Research

Research channel →