Just 30 RL Steps Drive Dramatic Model Gains With No Plateau in Sight

stochasticchasm · x · 2026-09-22

A paper headliner, stochasticchasm, shares that the most surprising finding was how much the model improved after only 30 RL steps — far fewer than expected — with no sign of a plateau, prompting the question of why the authors stopped training there.

Related event: Xiaomi MiMo: Possibly First 1M-Context RL Training, Big Gains in Just 30 Steps(4 posts)→

Original post →

More from Research

Research channel →