NVIDIA Details AI Factory Performance Per Watt
NVIDIA Blog · rss · 2026-07-14
NVIDIA emphasizes that performance per watt is the ultimate metric for AI infrastructure efficiency, directly dictating the revenue and profit of AI factories under fixed power budgets.
- Architectural Advantage: As frontier models shift to MoE architectures, the NVIDIA Blackwell NVL72 platform leverages hardware-software co-design to achieve far better performance per watt than Hopper on models like DeepSeek, GLM, and Kimi.
- Software Optimization: Combining technologies like NVFP4 quantization, disaggregated serving, and KV cache offloading significantly boosts single-GPU performance.
- Energy Management: The DSX MaxLPS platform dynamically allocates GPU power and supports liquid cooling, allowing operators to run up to 40% more GPUs under the same power budget.
- Production Validated: Companies like Anthropic, OpenAI, and CoreWeave have adopted this platform for large-scale inference services.
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11