NVL72+Groq Rack Throughput Analysis
Beth_Kindig · x · 2026-07-13
According to data cited by I/O Fund, if NVIDIA's Vera Rubin NVL72 rack is paired with a Groq 3 LPX rack:
- Token throughput per megawatt could increase by up to 35x
- For AI factories, the revenue opportunity could be 10x higher than Blackwell
- The cost of acquiring 1TB of HBM is over 5x that of using a Blackwell GPU plus 1TB of DDR5
The author suggests that as GPU scaling continues to be constrained by memory costs, offload engines will likely become a critical solution for scaling memory independently of GPUs in the next phase.
More from Infra
- Nvidia Rubin is coming, pointing to the next AI compute platform — ezyang · 2026-07-21
- Tesla’s FSD v14 Lite is reportedly headed to 4 million older HW3 cars — MatthewBerman · 2026-07-21
- TSMC’s 3nm utilization reportedly tops 120% as AI demand drives a $190B capex cycle — tengyanAI · 2026-07-21
- Nativ brings local AI model running to Mac with a desktop app and localhost API — Simon Willison · 2026-07-21
- Octen says agent search now runs at 62ms P50 with only a 6ms P90 gap — aakashgupta · 2026-07-21
- Zhipu acquires a compiler-team spinout to optimize AI inference on domestic chips — zephyr_z9 · 2026-07-21