NVIDIA Nemotron 3.5 Lightning Hits DeepInfra with 1M Token Context
gharik · x · 2026-08-11
DeepInfra has officially launched NVIDIA's Nemotron 3.5 Lightning model. Featuring a 30B MoE architecture with 3B active parameters, the model supports a massive 1 million token context window. It is touted as the fastest open model in its class, offering up to 4x higher throughput tailored for always-on agents.
Pricing is set at $0.05/M input tokens and $0.20/M output tokens, with full OpenAI API compatibility.
More from Infra
- UnslothAI Confirms Its Acceleration Tools Work Well with Apple's MLX Framework — danielhanchen · 2026-08-11
- Unsloth Desktop Launches: Open-Source App for Local LLM Training and Inference — danielhanchen · 2026-08-11
- Big Tech Capex Surge: Quarterly Bills to Exceed $10B Amidst AI Infrastructure Boom — Beth_Kindig · 2026-08-11
- PlayCanvas Launches Blazing Fast WebGPU 3DGS Renderer, Leading Mobile Performance — willeastcott · 2026-08-11
- Jensen Levels Up Neoclouds Against Hyperscalers With $500B Financing — firstadopter · 2026-08-11
- NVIDIA Launches Nemotron 3.5 Lightning and NeMo Switchyard for Agentic AI — nvidia · 2026-08-11