Classic MiMo architecture with huge RL batches: 1568 tasks x 16 rollouts per step, surprising results

stochasticchasm · x · 2026-09-22

The author notes MiMo uses a classic architecture — simple yet strong. The RL run used an enormous batch size at every step: 1568 tasks with 16 rollouts each, meaning the run consumed a lot of tokens. Even so, the results remain surprisingly good.

Related event: Xiaomi MiMo: Possibly First 1M-Context RL Training, Big Gains in Just 30 Steps(4 posts)→

Original post →

More from Research

Research channel →