Xiaomi MiMo: Possibly First 1M-Context RL Training, Big Gains in Just 30 Steps
Xiaomi's MiMo may be the first to run RL training with 1M context, using huge batches of 1568 tasks with 16 rollouts each. Researchers were surprised that just 30 RL steps yielded dramatic gains with no plateau in sight.
2026-09-22 ~ 2026-09-22 · 4 related posts
- Just 30 RL Steps Drive Dramatic Model Gains With No Plateau in Sight — stochasticchasm · 2026-09-22
- Classic MiMo architecture with huge RL batches: 1568 tasks x 16 rollouts per step, surprising results — stochasticchasm · 2026-09-22
- Xiaomi MiMo may be first RL at 1M context, with agent-as-a-judge as new scaling axis — stochasticchasm · 2026-09-22
1 near-duplicate retellings: stochasticchasm