NVIDIA Unveils Vera Rubin NVL72 Compute Tray
MoonL88537 · x · 2026-07-14
Citing an NVIDIA AI Infra briefing, the post highlights the compute power of the Vera Rubin NVL72 compute tray: a single tray delivers 200 AI petaFLOPs and can be assembled in just 1 minute.
It's noted as a single-width, 3rd-generation MGX rack solution featuring:
- No cables, no hoses, no fans
- 100% liquid-cooled, operating at 45°C
- Integrated Vera Rubin Superchips, ConnectX-9 SuperNICs, and BlueField-4 DPUs
- Designed for lower token costs and higher performance per watt
The original poster describes its massive scale as "adding 200 petaflops of global compute capacity every minute."
Related event: NVIDIA Unveils Vera Rubin NVL72 Interconnect Architecture(2 posts)→
More from Infra
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11