Xiaomi livestreams model training burning ~$10/second, mimo models in early Agentic RL stage
karminski3 · x · 2026-09-17
karminski3 spotted that Xiaomi is livestreaming its model training in real time, a rare public look at an ongoing training run, with several notable details:
- Two models in post-training: mimo-v2.6-pro and flash are being trained in parallel, with the page labeled RL — indicating the post-training (reinforcement learning) stage.
- It's Agentic RL: training progress is measured on the DeepSWE benchmark, but the run appears early-stage — mimo-v2.6-pro scores 62 on DeepSWE vs 74.2 for the leading DeepSeek-v4.1-flash.
- The burn rate: roughly $10+ per second. Cross-referencing page metrics — stage 10 took 58 minutes and generated 2.2B tokens — lets viewers estimate token throughput and back out GPU count and likely time-to-completion.
More from Infra
- Rust+Vulkan training backend runs 143 Transformer architectures without CUDA or PyTorch — PhysicsDisastrous462 · 2026-09-17
- Local Dual-RTX Pro Inference Rig: 150 tok/s Decode, 10K tok/s Prefill, Full Build Notes — No_Run8812 · 2026-09-17
- Open-Source Rust+Vulkan Training Backend Supports 143 Modern Transformer Architectures Without CUDA — PhysicsDisastrous462 · 2026-09-17
- GPU prices keep climbing as AI compute demand overwhelming supply — firstadopter · 2026-09-17
- Nearly all top-10 PFAS makers plan production hikes to serve AI chips and data center cooling — jathansadowski · 2026-09-17
- Idle models might be smarter: a musing on GEMV vs GEMM under low concurrency — karminski3 · 2026-09-17