NVIDIA Launches Nemotron 3.5 Lightning for Local AI Agents

ollama · x · 2026-08-12

NVIDIA announced that the Nemotron 3.5 Lightning model is now available on Ollama, capable of running entirely on local devices. Featuring a hybrid Mixture-of-Experts architecture with 30B total parameters and only 3B active per token, it is designed for agentic tasks requiring long-running context and tool calling.

The model supports up to 1M token context and optimized inference via multi-token prediction, offering up to 4x higher throughput. It is ideal for local personal assistants, coding sub-agents, and security operations, ensuring data privacy.

Related event: NVIDIA Launches Nemotron 3.5 Lightning Model and NeMo Switchyard Router(37 posts)→

Original post →

More from coding & agent

coding & agent channel →