DeepMind's guide to sharded matrix multiplication: the math behind training LLMs on 10k TPUs
zacharynado · x · 2026-09-16
Google DeepMind researchers (Jacob Austin, Roy Frostig, Reiner Pope et al.) published part 3 of 'How To Scale Your Model': a systematic theory of sharded matrix multiplication built on the cost of TPU communication primitives. It explains how to partition LLM parameters across thousands of accelerators efficiently, covering collective operations, partitioning notation, and why training and inference scale to larger topologies for different reasons (step time vs latency).
More from Infra
- RX 7900 XTX beats R9700 by 30% on GPT-OSS-20B local inference tests — glenbeer · 2026-09-16
- SemiAnalysis: Vera Rubin NVL144 hits ~7x tokens per MW vs Blackwell, over 2x profit per GW — sudoraohacker · 2026-09-16
- Musk explains why Terafab must exist: Taiwan chip risk plus capacity ceiling — elonmusk · 2026-09-16
- YuE2 music sampling only hits ~7 tokens/s on RX 9070 despite mostly idle VRAM — Prestigious-Kick7291 · 2026-09-16
- As LMs mature, the work reduces to two things: data and infra — saurabh_shah2 · 2026-09-16
- StartLux Interview: A 27B Local Model Rivals DeepSeek Giants via AutoResearch and Local RSI — 机器之心 · 2026-09-16