NVIDIA Launches Nemotron 3.5 Lightning: A 30B MoE Model Built for Always-On Agents

kuchaev · x · 2026-08-11

NVIDIA has introduced Nemotron 3.5 Lightning, a 30B hybrid Mamba-Transformer MoE model with 3B active parameters, distilled from Nemotron 3 Ultra.

Built specifically for "always-on" agents to handle high-volume, specialized tasks, it delivers up to 4x the output speed of similar-sized models. Key features include:

Inference framework SGLang has already announced Day 0 support for the model.

Related event: NVIDIA Open-Sources Nemotron Model and Agent Router(36 posts)→

Original post →

More from Infra

Infra channel →