Benchmark: Running vLLM with NVFP4 Weights on a Single GH200

llm_wizard · x · 2026-08-12

A developer shared inference benchmark results using a single GH200 (480GB) on Lambda Labs.

The test environment utilized vLLM 0.27.1 with NVFP4 weights enabled, forcing 256 output tokens at temperature 0, taking the median of 3 runs. The tweet also provided a link to download the official NVFP4 weights from NVIDIA AI.

Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→

Original post →

More from Infra

Infra channel →