Nemotron 3.5 Lightning Hits Baseten with 4x Throughput and Halved Costs

baseten · x · 2026-08-11

Inference platform Baseten announced day-0 support for NVIDIA's newly open-sourced Nemotron 3.5 Lightning model. Built for long-running agents, it is touted as the fastest open model in its class.

Compared to similar-sized open models, Nemotron 3.5 Lightning demonstrates significant advantages in production testing: delivering 4x higher throughput, 50% lower cost, and a 63.4% reduction in output tokens. The model utilizes a 30B MoE architecture (3B active parameters) and supports a 1M token context length.

Original post →

More from Models

Models channel →