Xiaomi's Livestreamed RL Training Burns ~$10/sec, Est. 6,000-8,000 H100s
karminski3 · x · 2026-09-17
Blogger karminski3 analyzed Xiaomi's livestreamed training run, concluding MiMo-v2.6-pro and flash are in the Agentic RL post-training phase, benchmarked on DeepSWE (Pro at 62 vs DeepSeek-v4.1-flash's 74.2 — still early).
His back-of-envelope math:
- Stage 10 generated 2.2B tokens in 58 minutes (630K tokens/sec) at 89K mean context length
- At 100-150 tps per H100, that implies 4,000-5,000 H100s for Pro decode plus 2,000-2,500 for flash — roughly 6,000-8,000 total
- Timing logs (58 min generation + 64 min backward pass ≈ 126 min/step) suggest the same GPUs flip between rollout inference and long-sequence backprop
- Over 60,000 concurrent Linux/Docker sandboxes (23,180 for Pro, 37,786 for flash) explain the $10/second burn rate
More from Infra
- Rust+Vulkan training backend runs 143 Transformer architectures without CUDA or PyTorch — PhysicsDisastrous462 · 2026-09-17
- Local Dual-RTX Pro Inference Rig: 150 tok/s Decode, 10K tok/s Prefill, Full Build Notes — No_Run8812 · 2026-09-17
- Open-Source Rust+Vulkan Training Backend Supports 143 Modern Transformer Architectures Without CUDA — PhysicsDisastrous462 · 2026-09-17
- GPU prices keep climbing as AI compute demand overwhelming supply — firstadopter · 2026-09-17
- Nearly all top-10 PFAS makers plan production hikes to serve AI chips and data center cooling — jathansadowski · 2026-09-17
- Idle models might be smarter: a musing on GEMV vs GEMM under low concurrency — karminski3 · 2026-09-17