Hands-on: NVIDIA Nemotron 3.5 Hits 250+ tokens/sec on 30B MoE
Arindam_1729 · x · 2026-08-11
A developer tested NVIDIA's newly released Nemotron 3.5 Lightning model. It is a 30B parameter MoE model (3B active parameters) designed specifically for high-volume, low-latency agentic tasks.
On the Nebius platform, actual inference speed reached an impressive 250-260 tokens/s. The tester noted that beyond its extreme speed, the model's reasoning performance is highly impressive, making it very promising for agentic workloads.
Related event: NVIDIA Launches Nemotron 3.5 Lightning Model and NeMo Switchyard Router(37 posts)→
More from Models
- Muse Glimmer 30B Hits 25 tok/s In-Browser on M4 Max via Custom WebGPU Kernels — xenovatech · 2026-08-12
- Anthropic Planted Thoughts Inside Claude's Neural Network to Test Introspective Awareness — johnmccrea · 2026-08-12
- Google Tops Text-to-Video Arena, Sora 2 Pro Misses Top 5 — arena · 2026-08-12
- LangChain Tests NVIDIA Switchyard: 93% of Agent Calls Handled by 30B Model, Cutting Costs 70% — LangChain · 2026-08-12
- Where Do Quantized Local LLMs Break? Reddit Users Share Experiences — d77chong · 2026-08-12
- Unreleased Anthropic Model Makes Surprising Progress on the Riemann Hypothesis — TechCrunch AI · 2026-08-12