NVIDIA Launches 30B MoE Model Optimized for High-Volume Agents

ziqiao_ma · x · 2026-08-12

NVIDIA has introduced Nemotron 3.5 Lightning, a new open 30B Mixture-of-Experts (MoE) model with only 3B active parameters. Optimized for throughput and latency, it is specifically built for always-on AI agents to handle high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models and is now available on the Tinker platform.

Related event: NVIDIA Launches Nemotron 3.5 Lightning Model and NeMo Switchyard Router(42 posts)→

Original post →

More from Models

Models channel →