NVIDIA and Baseten Optimize Nemotron 3 Ultra Inference, Boosting Concurrency 2.5x on Four B200 GPUs

NVIDIA and Baseten engineering teams detailed full-stack inference optimizations for Nemotron 3 Ultra NIM, covering caching, memory, parallelism and decoding. On four B200 GPUs, the optimized stack delivers up to 2.5x more concurrent users versus the unoptimized baseline.

2026-09-15 ~ 2026-09-15 · 2 related posts