NVIDIA Nemotron 3.5 Lightning Hits DeepInfra with 1M Token Context

gharik · x · 2026-08-11

DeepInfra has officially launched NVIDIA's Nemotron 3.5 Lightning model. Featuring a 30B MoE architecture with 3B active parameters, the model supports a massive 1 million token context window. It is touted as the fastest open model in its class, offering up to 4x higher throughput tailored for always-on agents.

Pricing is set at $0.05/M input tokens and $0.20/M output tokens, with full OpenAI API compatibility.

Original post →

More from Infra

Infra channel →