Nemotron 3.5 Lightning Hits Baseten with 4x Throughput and Halved Costs
baseten · x · 2026-08-11
Inference platform Baseten announced day-0 support for NVIDIA's newly open-sourced Nemotron 3.5 Lightning model. Built for long-running agents, it is touted as the fastest open model in its class.
Compared to similar-sized open models, Nemotron 3.5 Lightning demonstrates significant advantages in production testing: delivering 4x higher throughput, 50% lower cost, and a 63.4% reduction in output tokens. The model utilizes a 30B MoE architecture (3B active parameters) and supports a 1M token context length.
More from Models
- OpenRouter Data: Reasoning Model Token Share Exceeds 60% — Beth_Kindig · 2026-08-11
- Nvidia Releases Nemotron 3.5 Lightning Open 30B MoE Model — dr_alphalyrae · 2026-08-11
- Nemotron 3.5 Specs Revealed: 3.6B Active Params Matches Larger Models — thisguyknowsai · 2026-08-11
- LlamaIndex Launches ExtractBench: A New Benchmark for Enterprise Document Extraction — llama_index · 2026-08-11
- Benchmarking Meta's New Muse Glimmer: A 30B Model Optimized for Agents — altryne · 2026-08-11
- NVIDIA Updates 30B Hybrid Model: Distillation Delivers 70% Intelligence Boost — PavloMolchanov · 2026-08-11