Single GH200 Hits 4600 tok/s in Nemotron vLLM Test
A developer benchmarked the Nemotron 3.5 Lightning model on a single NVIDIA GH200 using vLLM with NVFP4 weights. The test achieved an impressive inference speed of over 4600 tok/s under 64 concurrent requests.
2026-08-12 ~ 2026-08-12 · 3 related posts
- Benchmark: Running vLLM with NVFP4 Weights on a Single GH200 — llm_wizard · 2026-08-12
- Nemotron 3.5 Hits 4,694 tok/s with 64 Concurrent Generations on a Single GH200 — pcuenq · 2026-08-12
1 near-duplicate retellings: aaditya_ai