Nemotron 3.5 Hits 4,694 tok/s with 64 Concurrent Generations on a Single GH200
pcuenq · x · 2026-08-12
A developer pushed the limits of concurrent generation with the Nemotron 3.5 Lightning model (30B-A3B NVFP4 checkpoint) on a single NVIDIA GH200.
When deployed using vLLM 0.27.1, single-stream inference reached 315 tok/s. Under a 64-way concurrent generation workload, aggregate throughput skyrocketed to 4,694 tok/s. This proves that launching dozens of AI agents simultaneously on a single GPU is now a reality, drastically lowering hardware barriers for high-concurrency agent applications.
Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→
More from Infra
- DeepSeek V4 Flash Hits 32.7 tok/s on AMD Strix Halo via 98GB GGUF — tensorqt · 2026-08-12
- NVIDIA Shares Guide on Running Nemotron 3 Ultra Locally on DGX Station — NVIDIAAI · 2026-08-12
- Phosphene Update: Run MiniMax H3 on 48GB Macs, 10s Single-Pass Renders — cocktailpeanut · 2026-08-12
- Understanding LLM inference: How Prefill and Decode phases impact speed — techNmak · 2026-08-12
- Running Muse Glimmer 30B on RX 7600 XT 16GB: Hits 20 t/s with Speculative Decoding — DanC403 · 2026-08-12
- vLLM Announces Day-0 Support for NVIDIA Nemotron 3.5 Lightning MoE Model — AccBalanced · 2026-08-12