Nemotron 3.5 Lightning Tested: 670 tokens/s Speedster for Agents
ArtificialAnlys · x · 2026-08-11
According to tests by Artificial Analysis, NVIDIA's newly released Nemotron 3.5 Lightning is an open-weights small model focused on extreme efficiency. With 31.6B total parameters (3.6B active), it achieves a dual leap in intelligence and speed while maintaining a compact footprint.
Key Data & Highlights:
- Major Intelligence Jump: Scores 24 on the Artificial Analysis Intelligence Index, a +9 point improvement over its predecessor (15), matching the larger gpt-oss-120b.
- Extreme Inference Speed: Achieves a blazing 670 tokens/second output speed on DeepInfra's NVFP4 quantized endpoint, vastly outperforming peers in its class.
- Agentic Step Change: Scores 24% on Terminal-Bench v2.1 (up from 7% previously) and surpasses the larger Nemotron 3 Super on GDPval-AA v2.
- Open Source & Availability: Released under the permissive OpenMDW-1.1 license for commercial use, features a 1M token context window, and is available across multiple inference providers.
Related event: Nemotron 3.5 Lightning Tested for High-Speed Inference(2 posts)→
More from Models
- UnslothAI Confirms Its Acceleration Tools Work Well with Apple's MLX Framework — danielhanchen · 2026-08-11
- How to Strip Claude's Text Watermark? Users Test Translation Workarounds — churchkey · 2026-08-11
- Anthropic Criticized for AI Watermark Strategy That Could Drive Users Away — Brian821 · 2026-08-11
- Daybreak Blue in Codex Clarified: Not GPT-5.6 Cyber, But Security-Tailored GPT-5.6 Sol — Angaisb_ · 2026-08-11
- Anthropic to Embed Imperceptible Watermarks in Claude Text for EU Compliance — lilyraynyc · 2026-08-11
- NVIDIA releases Nemotron 3.5 Lightning: 30B LatentMoE, 3B active, 1M context — huggingface · 2026-08-11