Hands-on: NVIDIA Nemotron 3.5 Hits 250+ tokens/sec on 30B MoE

Arindam_1729 · x · 2026-08-11

A developer tested NVIDIA's newly released Nemotron 3.5 Lightning model. It is a 30B parameter MoE model (3B active parameters) designed specifically for high-volume, low-latency agentic tasks.

On the Nebius platform, actual inference speed reached an impressive 250-260 tokens/s. The tester noted that beyond its extreme speed, the model's reasoning performance is highly impressive, making it very promising for agentic workloads.

Related event: NVIDIA Launches Nemotron 3.5 Lightning Model and NeMo Switchyard Router(37 posts)→

Original post →

More from Models

Models channel →