~$11 per billion tokens: decentralized-trained model's inference optimizations revealed
bittingthembits · x · 2026-09-13
jondurbin quotes his team's decentralized training run, priced around $11/b tokens with impressive MFU. But he stresses the real point isn't the training itself — it's the inference optimizations the architecture enables:
- sparse fp4 native compute for routed experts
- fixed KV cache for most attention
- sparse latent for the rest, tiny model weights
Numbers come from an unoptimized vllm branch on their own arch, and he caveats the comparison isn't fully fair. Core claim: limitless tokens on cheap hardware, using a few nodes from liumio.
More from Infra
- SGLang Team Open-Sources Miles v0.1: 744B Model RL on 64 GPUs at 263s/Step — aigclink · 2026-09-13
- Reading up on terahertz pulse reflectometry for inspecting CoWoS packaging layers — jwt0625 · 2026-09-13
- DeepMind Chief Strategist: AI Infrastructure Spending Is a Bet on Recursive Self-Improvement — rohanpaul_ai · 2026-09-13
- AI valuations can't all be right: memory at 3-5x PE vs premium infrastructure, says Gavin Baker — rohanpaul_ai · 2026-09-13
- Draw Things launches Local Code beta: local coding agents on Mac at ~980 tok/s prefill — liuliu · 2026-09-13
- Macrocosmos launches IOTA for liquid training on scattered, disaggregated compute — markjeffrey · 2026-09-13