FULL STORY
Xiaomi MiMo's Million-Token RL Training Details Spark Debate
Xiaomi's MiMo sparked buzz as reportedly the first to run RL training on 1M-token contexts, followed by discussion of its paper on stabilizing GRPO for MoE models.
2026-09-22 ~ 2026-09-22 · 2 episodes · 15 posts
Episode 1 · Xiaomi MiMo's 1M-Context RL Training Details Spark Debate (2026-09-22, 13 posts)
Details of Xiaomi MiMo's RL training have sparked heated discussion overseas. stochasticchasm believes it may be the industry's first reinforcement learning training run on a 1 million token context, achieving dramatic performance gains in only about 30 steps with no plateau visible in the training curve — and Agent-as-judge may become a new scaling axis.
Confirmed
- MiMo uses a classic architecture, described by commenters as simple and strong
- Training batches are enormous per step: 1568 tasks × 16 rollouts per task, consuming massive token counts per round
- stochasticchasm notes the model achieved dramatic gains after only 30 RL steps, far fewer than he expected
- There are no signs of a plateau in the training curve; both stochasticchasm and commenter nrehiew questioned why training wasn't continued
Why it matters
- If 1M-context RL training is confirmed, it would be a breakthrough on a new scaling axis, potentially benefiting long-context Agent tasks
- Reaching near-frontier performance in 30 steps suggests the sample efficiency of very-large-batch RL may far exceed prior assumptions, offering useful guidance for future training strategies
- Just 30 RL Steps Drive Dramatic Model Gains With No Plateau in Sight — stochasticchasm · 2026-09-22
- Model improved dramatically in just 30 RL steps with huge batch sizes — stochasticchasm · 2026-09-22
- Classic MiMo architecture with huge RL batches: 1568 tasks x 16 rollouts per step, surprising results — stochasticchasm · 2026-09-22
- Xiaomi MiMo may be first RL at 1M context, with agent-as-a-judge as new scaling axis — stochasticchasm · 2026-09-22
- Researchers debate GRPO at scale: standard formulation, missing batch-size ablations — stochasticchasm · 2026-09-22
- Standard GRPO at 1M scale: why no critic models, and what counts as "behaviors"? — stochasticchasm · 2026-09-22
- RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization — stochasticchasm · 2026-09-22
- RL training detail: re-prefilling instead of PipelineRL's cached KV, with batch size framed as a GPU-utilization lever — stochasticchasm · 2026-09-22
- RL run spends a lot of compute on graders; is re-prefilling worth it over 30 steps — stochasticchasm · 2026-09-22
- Xiaomi's model jumps near the frontier with just 30 RL steps, no plateau in sight — teortaxesTex · 2026-09-22
- 13% Policy Drift Staleness Suggests RL Envs Needed No Curriculum, Tokenbender Argues — tokenbender · 2026-09-22
- RL training stability tricks: prompt-mean loss, asymmetric clipping and entropy-adaptive GRPO — tokenbender · 2026-09-22
- Reading Xiaomi's MoE RL Release: Specialist Teachers Beat Big Mixed RL on Hard Tasks — tokenbender · 2026-09-22
Episode 2 · Xiaomi's MoE RL stability paper: router-aware GRPO tweaks explained (2026-09-22, 2 posts)
tokenbender breaks down Xiaomi's arXiv paper on stable MoE RL: rather than replacing GRPO, the team adds targeted fixes such as router-aware importance-sampling rescaling and prompt-mean loss for stability and scale.
- Sticking with GRPO: The Four Targeted Mods Behind Xiaomi's Stable MoE RL — tokenbender · 2026-09-22
- New arXiv Paper: Router-Aware Importance Sampling Stabilizes MoE RL Training — tokenbender · 2026-09-22