Standard GRPO at 1M scale: why no critic models, and what counts as "behaviors"?
stochasticchasm · x · 2026-09-22
stochasticchasm analyzes a RL training setup run at 1M scale: it uses a very standard GRPO formulation and doesn't yet adopt the critic models some teams have started using. He speculates large batches make gradients noisy enough-free without scaling group size, and wonders what "behaviors" means precisely — perhaps relative rankings of optimized code/performance — hoping the released environments include more detail.
Related event: Debate Erupts Over Million-Scale GRPO RL Training Details(3 posts)→
More from Research
- New Paper: Test-Time Communication Makes Agent Teams Beat Independent Agents — DimitrisPapail · 2026-09-22
- New Paper: Test-Time Communication Makes Agent Teams Beat Independent Agents — DimitrisPapail · 2026-09-22
- Mathematicians clash over formal proofs: who voted to change math's rules? — jessi_cata · 2026-09-22
- Lean vs ZFC: the rules of mathematical proof weren't changed by any vote — jessi_cata · 2026-09-22
- David Krueger: four unresolved foundational problems stand between us and safe AI — DavidSKrueger · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22