SemiAnalysis: NVIDIA Rubin with vLLM Delivers 3.2x Profit per Gigawatt, up to 10x Perf per Dollar vs GB300 NVL72
woosuk_k · x · 2026-10-10
- SemiAnalysis reports that NVIDIA's next-gen Rubin platform running on the widely used production inference engine vLLM delivers 3.2x better profit per gigawatt and up to 10x better performance per dollar than even GB300 NVL72.
- Because the benchmark uses vLLM, a mainstream production LLM engine, the claim reflects real-world deployment economics rather than a vendor-curated benchmark.
- If Rubin ships as claimed, it could significantly shift per-unit inference economics in a market where power and cost dominate.
More from Infra
- Full AI video made at home on a single RTX 3090 with Qwen, FLUX, MiniMax and YuE2 — solyarisoftware · 2026-10-10
- Baseten launches Project Beacon with Goodfire for inline open-model safety — baseten · 2026-10-10
- Huawei launches OceanStor M900 context memory storage for hyperscale AI data centers — pstAsiatech · 2026-10-10
- Morgan Stanley: Amazon Trainium3/4 to drive Intel EMIB-M packaging, capacity scaling 8x by 2028 — pstAsiatech · 2026-10-10
- Data centers could become SpaceX's biggest revenue stream, with $1T revenue on the horizon — Dr_Singularity · 2026-10-10
- New gTLD application for .lan shows why internal services need real registered domains — evilsocket · 2026-10-10