Deep Dive: TPU and GPU Cluster Communication
thehiphopswami · x · 2026-07-15
This is an in-depth blog post about TPU/GPU clusters and collective communication, focusing on the underlying mechanics of scaling training and inference.
Topics covered include:
- TPU cluster topology: superpods, slices, DCN, PCIe, ICI
- How All-Gather, Reduce-Scatter, and All-Reduce work
- The role of All-to-All in MoE token dispatch
- NVIDIA GPU cluster topologies and related communication paths
The author's goal is to demystify these seemingly abstract communication primitives, helping readers truly grasp the foundations of scaling MoE and dense transformer models.
Related event: Deep Dive on Collective Communication in TPU/GPU Clusters Gains Traction(5 posts)→
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11