NVIDIA ships Nemotron 3.5 Lightning, a 30B MoE built for agentic execution at 4x speed

cwolferesearch · x · 2026-08-18

NVIDIA's expanding Nemotron family now includes Nemotron 3.5 Lightning, a small MoE model (30B params, 3B active) targeting high-throughput, low-latency use cases.

Its role is the execution layer of agentic workflows: frontier reasoning models like Nemotron 3 Ultra handle planning and orchestration, while 3.5 Lightning handles high-volume steps like validating tool outputs, executing commands, and solving simpler subtasks. It is specifically trained for agentic use and compatible with frameworks like hermes and openclaw.

For inference efficiency, it is trained with multi-token prediction for speculative decoding and ships drafters for DSpark/DFlash plus quantized checkpoints. It runs 4x faster than similar-size models and scores notably well on the Artificial Analysis intelligence index for its speed category. NVIDIA also provides out-of-the-box finetuning recipes for post-training reliable, low-cost models.

Original post →

More from Infra

Infra channel →