Nemotron 3.5 Hits 4,694 tok/s with 64 Concurrent Generations on a Single GH200

pcuenq · x · 2026-08-12

A developer pushed the limits of concurrent generation with the Nemotron 3.5 Lightning model (30B-A3B NVFP4 checkpoint) on a single NVIDIA GH200.

When deployed using vLLM 0.27.1, single-stream inference reached 315 tok/s. Under a 64-way concurrent generation workload, aggregate throughput skyrocketed to 4,694 tok/s. This proves that launching dozens of AI agents simultaneously on a single GPU is now a reality, drastically lowering hardware barriers for high-concurrency agent applications.

Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→

Original post →

More from Infra

Infra channel →