100B Model Trained Across 5 Data Centers on Plain Internet Links at 30.8% MFU
markjeffrey · x · 2026-09-25
- IOTA (Macrocosmos, on Bittensor) trained a 100B parameter model across five data centers using single A100s connected over ordinary internet links.
- Results: 30.8% average MFU, 38% sustained peak, 65% the speed of the same job on co-located high-bandwidth hardware — reportedly the highest MFU reported for a distributed pipeline-parallel run.
- The run processed 1.1B tokens over two days and stopped on cost grounds; the goal was proving a model that size holds together on disaggregated compute. It did.
- IOTA positions itself as liquid/disaggregated training infrastructure that unifies scattered, underused compute into usable training capacity.
More from Infra
- AI energy startup Parallax launches with $117m from Founders Fund, Lux, Greylock and others — graceisford · 2026-09-25
- AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026 — PyTorch · 2026-09-25
- Nebius/WEKA benchmark: shared KV cache lifts agentic inference throughput 2.4x with 93% hit rate — AccBalanced · 2026-09-25
- Burkov's TP Weekly #179: GPU rent vs buy, llm-d serving 753B model at 5-10x lower cost — burkov · 2026-09-25
- YuE2 music model with full CoT runs on iPhone in just 1.7GB of memory — Acceptable-Cycle4645 · 2026-09-25
- Microsoft Fabric's Efficient Scaledown Cuts Shuffle-Heavy Spark Costs by 43% — adnan_hashmi · 2026-09-25