NVIDIA Walks Through Deploying Nemotron 3.5 Lightning on DGX Spark for Agents
NVIDIA Developer · youtube · 2026-08-11
NVIDIA has released a technical walkthrough demonstrating how to deploy the Nemotron 3.5 Lightning model on DGX Spark for high-volume agentic workloads.
- Model Specs: Nemotron 3.5 Lightning is a 30B parameter MoE model with 3B active parameters, designed specifically as a fast execution model for always-on agent systems. It supports up to 1M tokens of context.
- Deployment: The tutorial covers deployment using vLLM and DSpark speculative decoding, connecting the model to agent harnesses like OpenCode.
- Performance: On a single DGX Spark, the setup can handle 16 concurrent agent tasks, generating over 500 tokens per second.
Related event: NVIDIA Open-Sources Nemotron Model and Agent Router(37 posts)→
More from coding & agent
- Payments and Service Discovery for Autonomous Agents: No Standard Yet — Nata_Elisym · 2026-08-12
- Agent Success Rate Drops on Repeat: Paper Reveals Computer Use Reliability Trap — xwang_lk · 2026-08-12
- Cursor reportedly set to launch Composer 3 as users await evals — kimmonismus · 2026-08-12
- Developer Showcases x40-Powered Agent Endpoint for Reverse Phone Lookup — MurrLincoln · 2026-08-12
- Fine-tuned 30B NVIDIA Model Beats 50x Larger Models in Code Verification — fhuszar · 2026-08-12
- Coldtea launches ADE for self-driving software, integrating coding agents and monitoring — EXM7777 · 2026-08-12