NVIDIA Launches Nemotron 3.5 Lightning: 30B Params with 3B Active, Built for Agent Execution
heyshrutimishra · x · 2026-08-11
NVIDIA released Nemotron 3.5 Lightning, a model optimized for the execution layer of AI agents.
- Core Architecture: 30B parameter MoE model with only 3B active parameters per token, delivering 30B-level capacity at a fraction of the compute cost.
- Design Philosophy: Agents spend 90% of their time on routine tasks like tool calls and result validation. Lightning handles these tasks cheaply, leaving frontier models to handle planning.
- Performance: Achieves 86% accuracy on PinchBench; completes 10,000 agent tasks 30% faster than Qwen3.6 35B at similar accuracy.
- Deployment: Features baked-in speculative decoding and NVFP4 quantization, allowing the same checkpoint to run across Blackwell, Hopper, and Ampere architectures.
- Open Source: Fully open under OpenMDW-1.1, including weights, data, and recipes.
Related event: NVIDIA Open-Sources Nemotron Model and Agent Router(37 posts)→
More from coding & agent
- Agent Success Rate Drops on Repeat: Paper Reveals Computer Use Reliability Trap — xwang_lk · 2026-08-12
- Cursor reportedly set to launch Composer 3 as users await evals — kimmonismus · 2026-08-12
- Developer Showcases x40-Powered Agent Endpoint for Reverse Phone Lookup — MurrLincoln · 2026-08-12
- Fine-tuned 30B NVIDIA Model Beats 50x Larger Models in Code Verification — fhuszar · 2026-08-12
- Coldtea launches ADE for self-driving software, integrating coding agents and monitoring — EXM7777 · 2026-08-12
- Post-Trained NVIDIA Nemotron Beats Claude Opus in Legal Agent Tasks — ctnzr · 2026-08-12