Nvidia's Cluster Scale Metrics Spark Debate
suchenzang · x · 2026-07-17
This comment discusses how Nvidia's cluster scale is described:
- Clarifying what "scale up domain" means: it could refer to the connection range for a single set of TP/PP model weights, or the maximum size of a super pod that can be linked without performance degradation.
- It also notes a highly unusual unit: 72 of which Nvidia, questioning whether this is a typical server-side Nvidia cluster description.
- The provided numbers show a scale up of 8192 chips versus Nvidia's 72; the maximum scale out size is 500k.
The core takeaway is that different vendors and systems have vastly different metrics and limits regarding large-scale cluster interconnects and scaling.
Related event: Huawei 950 SuperPoD Sparks Debate Over Cluster Scale(9 posts)→
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11