NVIDIA Nemotron 3.5 Lightning Goes Live on CoreWeave Serverless

wandb · x · 2026-08-12

CoreWeave announces that the NVIDIA Nemotron 3.5 Lightning model is now live on its serverless inference platform. Distilled from Nemotron 3 Ultra, this customized 30B MoE model (3B active parameters) is designed specifically for always-on agents.

The company highlights that it is the fastest open model in its class, ideal for high-volume tasks like coding, tool calling, and multi-turn agentic workflows, allowing developers to deploy without managing underlying infrastructure.

Related event: Nemotron 3.5 Lightning Hits Baseten with 4x Throughput and Halved Costs(2 posts)→

Original post →

More from Infra

Infra channel →