Xiaomi MiMo runs full-parameter RL on 310B model across 1000+ TPUs with JAX
AccBalanced · x · 2026-09-22
The Xiaomi MiMo team details scaling RL with JAX + TPU: full-parameter RL on the 310B MiMo-V2.6 and 1000+ stable training steps across 1000+ TPUs. Highlights: scaling up is mostly a config change rather than a rewrite; optimized vLLM inference speeds up rollouts; bitwise trainer–sampler agreement in validation; trainer and sampler share one TPU ICI fabric, transferring all 310B parameters in under 2 seconds. Built with Berkeley AI and Google Cloud, more details promised.
Related event: Xiaomi Releases MiMo-V2.6: Scaling RL with JAX+TPU at 310B Scale(10 posts)→
More from Infra
- crabbox now runs on boxd: isolated KVM microVMs with ms boot times for repo commands — steipete · 2026-09-23
- AI data centers now demand up to 1,000 MW, a 200x jump in power — Olivier__OG · 2026-09-23
- Rust weight-streaming engine runs FLUX.2 9B on a 12GB RTX 3060, ~9% faster with 2GB resident pool — madtune22 · 2026-09-23
- AMD MI355X Hits 1.7x Perf-Per-Dollar vs DGX B300 via SGLang KV Cache Fix — zephyr_z9 · 2026-09-23
- Apsara Conference's real story: China's full AI stack, chips to robotics — manishkhosiya · 2026-09-23
- Local Qwen3.8-27B Runs Typed Decisions in <10GB VRAM at 170ms — kyr0x0 · 2026-09-23