NVIDIA Launches Nemotron 3.5 Lightning: A 30B MoE Model Built for Agents

baseten · x · 2026-08-11

NVIDIA has released the Nemotron 3.5 Lightning model, now available on the Baseten platform. Distilled from the larger Nemotron 3 Ultra, this model is specifically designed for long-running agentic workflows. It features a 30B Mixture-of-Experts (MoE) architecture with only 3B active parameters and supports a 1 million token context length.

According to official and production testing data, the model delivers 4x higher throughput, 50% lower costs, and 63.4% fewer output tokens compared to similar-sized open models. Additionally, the CodeRabbit team fine-tuned this model using Baseten Training, achieving an 80.4% routing agreement, a 4.6% increase over the 75.8% baseline.

Related event: NVIDIA Open-Sources Nemotron 3.5 Lightning Model and NeMo Switchyard(24 posts)→

Original post →

More from coding & agent

coding & agent channel →