Mixed 180 Consumer GPUs Train 50B Tokens: 100 Interruptions Add Only 2% Cost
bittingthembits · x · 2026-08-07
MacrocosmosAI's team IOTA SN9 project demonstrates distributed training capabilities in a highly unstable environment. The experiment clustered over 180 mixed RTX 4090/5090 GPUs over the ordinary internet.
- Performance: Achieves 65% of equivalent datacenter speed at a fraction of the cost.
- Fault Tolerance: In a 50B token training run, 100 node interruptions or departures resulted in only a 2% additional token cost.
- Technical Approach: By compressing and reducing the data that needs to be moved between GPUs, unreliable global compute acts as one reliable AI datacenter.
More from Infra
- SpaceX's New Data Center to Feature On-site Natural Gas and Battery Arrays — BenBajarin · 2026-08-07
- Generating a 3-Minute AI Music Video Locally on a Single RTX 3090 with MiniMax T2V — Inevitable_Emu2722 · 2026-08-07
- 38% of Americans Live Within 5 Miles of a Data Center — neil_chilson · 2026-08-07
- Useful Forge Neo Extensions: Convert to Int8 in Seconds with No Visible Quality Loss — cradledust · 2026-08-07
- AI Agent Inference Costs Drop: Running Hot for a Month is Now Cheaper Than a Burrito — intellectronica · 2026-08-07
- Speculative Decoding Will Reshape LLM API Training Terms of Service — charles_irl · 2026-08-07