ezyang's DeepSeek-V3 series part 3: roofline analysis for training DSv3 on Hopper
ezyang · x · 2026-10-03
PyTorch core developer Edward Z. Yang published part 3 of his DeepSeek-V3 training series, applying roofline analysis to reproduce DSv3's configuration on a H800 cluster.
Key points:
- Earlier parts found a parallelism/activation-checkpointing/low-precision setup that fits DSv3 on 2048 H800 GPUs, closely matching DeepSeek's own choices — expected, since optimizing a fixed architecture on the hardware it was designed for converges to the same config
- Roofline analysis: with FLOPs fixed, system performance hinges on distributed communication, analogous to kernel compute/bandwidth boundness
- Key conclusion: FSDP (ZeRO-3) is a bad fit for DSv3 — it saturates InfiniBand bandwidth — so you can rule it out analytically without building both systems and benchmarking at full cluster scale
- Also covers speed-of-light bound estimation and previews infra/architecture codesign topics like expert sparsity
More from Infra
- Cohere and vLLM co-host Toronto meetup on open weights and inference — cohere · 2026-10-03
- 17 configs, 3 weeks: Qwen3.8 low-latency cookbook hits 2.8x with DFlash2 on 27B — FantasticNature7590 · 2026-10-03
- TSMC evaluating multi-billion dollar Texas fab campus amid surging US demand from Nvidia, Apple — Beth_Kindig · 2026-10-03
- Discrete diffusion delivers provably lossless LLM inference speedups, drop-in for training — Cohere · 2026-10-03
- Runner pays $4,500 for a third GPU to run near-frontier models locally — marian_nmt · 2026-10-03
- 90% of AI chip value accrues to incumbents, no OpenAI equivalent for years — menhguin · 2026-10-03