Model improved dramatically in just 30 RL steps with huge batch sizes

stochasticchasm · x · 2026-09-22

stochasticchasm is surprised a model improved so much in only 30 RL steps — far fewer than expected. Each step had a massive batch of 1,568 tasks × 16 rollouts, consuming many tokens, yet there's no sign of a plateau, raising the question of why the run wasn't continued.

Related event: Xiaomi MiMo: Possibly First 1M-Context RL Training, Big Gains in Just 30 Steps(4 posts)→

Original post →

More from Research

Research channel →