NVIDIA and Baseten Optimize Nemotron 3 Ultra Inference, Boosting Concurrency 2.5x on Four B200 GPUs
NVIDIA and Baseten engineering teams detailed full-stack inference optimizations for Nemotron 3 Ultra NIM, covering caching, memory, parallelism and decoding. On four B200 GPUs, the optimized stack delivers up to 2.5x more concurrent users versus the unoptimized baseline.
2026-09-15 ~ 2026-09-15 · 2 related posts
- NVIDIA: full-stack NIM tuning delivers 2.5x more concurrent users on Nemotron 3 Ultra — NVIDIAAI · 2026-09-15
- Baseten tunes Nemotron 3 Ultra NIM serving, 2.5x more concurrent users on 4x B200 — baseten · 2026-09-15