Single GH200 Hits 4600 tok/s in Nemotron vLLM Test

A developer benchmarked the Nemotron 3.5 Lightning model on a single NVIDIA GH200 using vLLM with NVFP4 weights. The test achieved an impressive inference speed of over 4600 tok/s under 64 concurrent requests.

2026-08-12 ~ 2026-08-12 · 3 related posts

1 near-duplicate retellings: aaditya_ai