Xiaomi's MiMo V2.6 RL training livestream burns $3.24M in 4.5 days
teortaxesTex · x · 2026-09-20
Xiaomi publicly livestreamed the RL training of MiMo-V2.6 Pro and Flash, racking up roughly $3.24M in compute costs in 4.5 days ($2.39M Pro + $0.85M Flash).
- Pro still training at step 29 ($20.6K/hr, 3.0B tokens/step), DeepSWE 70.92
- Flash stalled over a day at step 30 ($10.3K/hr, 3.7B tokens/step), DeepSWE 65.68
- Setup: 1568 prompts × 16 async rollouts, 23 harnesses mixed, 60% coding data
- Hiccups included GPU OOM, VRAM limits, grader outages, and an infra failure hidden for 3 hours
The poster hopes other companies pull the same stunt.
More from Infra
- Google's Agent Substrate detailed: AX app layer on managed agentic compute infra — rakyll · 2026-09-21
- Why sandbox-as-a-service startups are booming — and whether labs will just build it themselves — dejavucoder · 2026-09-21
- Agents may discover million-times-cheaper training, making data centers look silly, predicts Steve Moraco — menhguin · 2026-09-21
- llama.cpp PR enables sparse FlashAttention for Qwen, another inference speedup — jacek2023 · 2026-09-21
- Analyst: Agentic CPU-to-GPU Ratio Won't Hit 40:1, More Like 4:1 — BenBajarin · 2026-09-20
- QwenImage 2.1 INT4 runs in just 4GB VRAM with ConvRot quantization — reeight · 2026-09-20