NVIDIA Launches Nemotron 3.5 Lightning: A 30B MoE Model Built for Always-On Agents
kuchaev · x · 2026-08-11
NVIDIA has introduced Nemotron 3.5 Lightning, a 30B hybrid Mamba-Transformer MoE model with 3B active parameters, distilled from Nemotron 3 Ultra.
Built specifically for "always-on" agents to handle high-volume, specialized tasks, it delivers up to 4x the output speed of similar-sized models. Key features include:
- Three built-in speculators: MTP, DFlash, and DSpark
- Trained for agent harnesses including coding, tool use, and multi-turn interactions
- Up to 1M-token context length, with BF16 and NVFP4 precision support
- Runs across Jetson, DGX Spark, Hopper, and Blackwell architectures
Inference framework SGLang has already announced Day 0 support for the model.
Related event: NVIDIA Open-Sources Nemotron Model and Agent Router(36 posts)→
More from Infra
- Bitcoin miners pivot to AI: ROIC ~3x mining, 70-75% revenue expected from AI — bittingthembits · 2026-08-12
- Single Used GPU Matches Claude 3 Opus: Local AI Potential Underestimated — DynamicWebPaige · 2026-08-12
- NVIDIA Nemotron 3.5 Lightning Goes Live on CoreWeave Serverless — wandb · 2026-08-12
- AWS Releases Reference Architecture for Enterprise Claude Apps Gateway — AWS ML Blog · 2026-08-11
- NVIDIA Expert: Multi-Token Techniques Become Day-Zero Norm for Inference — PavloMolchanov · 2026-08-11
- Autonomous Computer: $26K Dual RTX 5090 Workstation Targets Local Frontier Models — dee_hw · 2026-08-11