AI Infra Summit: Vera Rubin + DSX push tokens per megawatt, up to 35x vs GB200

NVIDIA Blog · rss · 2026-09-16

At AI Infra Summit (8,000+ attendees), NVIDIA's Ian Buck announced: Amazon Annapurna Labs co-developing NVHBM custom HBM; d-Matrix joining NVLink Fusion; Lambda validating DSX MaxLPS with 24% more token throughput in the same power budget; Vera Rubin NVL72 + Groq 3 LPX delivering up to 35x tokens/MW vs GB200 NVL72 for long-context 2T+ parameter models (2,529 output tokens/s/user on 100K-context workloads); and SemiAnalysis AgentX results showing up to 30x throughput/MW and 45x lower cost per million tokens versus GB300 NVL72. Emerald AI also demoed AI factories as dispatchable grid resources with Silicon Valley Power.

Original post →

More from Infra

Infra channel →