NVIDIA Unveils Rubin NVL72 Interconnect Architecture
AravSrinivas · x · 2026-07-14
NVIDIA detailed the scale-up interconnect architecture of the Vera Rubin NVL72, centered around the 6th-gen NVLink. It uses NVLink 6 Switch trays to form a fully connected system of 72 Rubin GPUs, transmitting data via a vertical NVLink spine. NVIDIA emphasized that this design aims to lower token costs and deliver 10x better performance per watt compared to the previous generation, thanks to enhanced reliability and shorter interconnect paths.
The post also mentioned a single-width MGX rack design: integrating the horizontal NVLink scale-up traverse directly onto the PCB to reduce cable complexity and improve deployment stability at scale.
Related event: NVIDIA Unveils Vera Rubin NVL72 Interconnect Architecture(2 posts)→
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11