MiMo RL livestream: infra restarts waste 33% of compute, over $400K lost
dustinvtran · x · 2026-09-19
Day 3 of the public MiMo RL training livestream revealed brutal infra economics: task restarts caused by infra issues (OOM, network, VRAM, silent data-induced failures) cost the Pro model 33% of effective compute — over $400K, with each restart losing tens of thousands of dollars. Former xAI/Gemini engineer dustinvtran amplified it: captaining an RL run that actually ships means discovering and patching infra and recipe issues mid-run; if you think there's no issue, you don't have enough metrics or aren't spot-checking enough traces.
More from Infra
- Dev builds daily-driver AI coding environment on AWS Lambda MicroVMs with S3 persistence — blaizedsouza · 2026-09-19
- Running 35B–284B MoEs at 128k Context on a 12GB A2000: Measured, Plus What Didn't Matter — HomoAgens1 · 2026-09-19
- Joseph Suarez hits SOTA in 5 hours on consumer cards in a warehouse — yacineMTB · 2026-09-19
- Why Jev-class models could become a near-free judgment primitive running on-device — signulll · 2026-09-19
- Pedro Domingos: Opposing Data Centers Means Keeping Your Country Stupid — pmddomingos · 2026-09-19
- Pedro Domingos: OpenAI and Anthropic's real moat is their massive secured compute — pmddomingos · 2026-09-19