Xiaomi MiMo runs full-parameter RL on 310B model across 1000+ TPUs with JAX

AccBalanced · x · 2026-09-22

The Xiaomi MiMo team details scaling RL with JAX + TPU: full-parameter RL on the 310B MiMo-V2.6 and 1000+ stable training steps across 1000+ TPUs. Highlights: scaling up is mostly a config change rather than a rewrite; optimized vLLM inference speeds up rollouts; bitwise trainer–sampler agreement in validation; trainer and sampler share one TPU ICI fabric, transferring all 310B parameters in under 2 seconds. Built with Berkeley AI and Google Cloud, more details promised.

Related event: Xiaomi Releases MiMo-V2.6: Scaling RL with JAX+TPU at 310B Scale(10 posts)→

Original post →

More from Infra

Infra channel →