Single GH200 NVFP4 Inference Benchmark: vLLM Performance Baselines Shared

aaditya_ai · x · 2026-08-12

A developer shared inference performance baseline results running on a single NVIDIA GH200 480GB.

Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→

Original post →

More from Infra

Infra channel →