Benchmark: Running vLLM with NVFP4 Weights on a Single GH200
llm_wizard · x · 2026-08-12
A developer shared inference benchmark results using a single GH200 (480GB) on Lambda Labs.
The test environment utilized vLLM 0.27.1 with NVFP4 weights enabled, forcing 256 output tokens at temperature 0, taking the median of 3 runs. The tweet also provided a link to download the official NVFP4 weights from NVIDIA AI.
Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→
More from Infra
- v100-skinny: Open Source Kernels Push Qwen 27B Inference to 366 t/s on V100 GPUs — Simple_Library_2700 · 2026-08-12
- CoreWeav Adds Over $2.5B in New Customer Commitments in Early Q3 — firstadopter · 2026-08-12
- WeAreDevs Talk: Providing On-Demand Compute for AI Agents — steren · 2026-08-12
- Hetzner Launches Experimental Free LLM Inference API Featuring DeepSeek and More — AccBalanced · 2026-08-12
- Transformers.js Surpasses 10 Million Monthly Downloads, Rapid Growth Continues — nicodotdev · 2026-08-12
- d-Matrix Chip Claims 20x Speedup for Qwen Inference — TheKanter · 2026-08-12