NVIDIA Nemotron 3.5 Lightning Goes Live on CoreWeave Serverless
wandb · x · 2026-08-12
CoreWeave announces that the NVIDIA Nemotron 3.5 Lightning model is now live on its serverless inference platform. Distilled from Nemotron 3 Ultra, this customized 30B MoE model (3B active parameters) is designed specifically for always-on agents.
The company highlights that it is the fastest open model in its class, ideal for high-volume tasks like coding, tool calling, and multi-turn agentic workflows, allowing developers to deploy without managing underlying infrastructure.
Related event: Nemotron 3.5 Lightning Hits Baseten with 4x Throughput and Halved Costs(2 posts)→
More from Infra
- Mojo 1.0 Released: The Systems Language for the AI Era — clattner_llvm · 2026-08-12
- Nvidia's Switchyard Router Reshuffles AI Models Mid-Task, Cutting Costs to 1/3 — CackleRooster · 2026-08-12
- Data Center Tax Boom Leads to 10 Years of Property Tax Cuts in Virginia — robleclerc · 2026-08-12
- Breaking VM Barriers: Apple Silicon LLM Inference Runs 16x Faster — petrusenko_max · 2026-08-12
- Ling-3.0-flash Quantization Benchmarks: MoE Architecture Preserves Decode Speed — AcanthisittaOk1699 · 2026-08-12
- SD Video Optimization: CK Cuts Generation Time to 473s, but Degrades Prompt Adherence — switch2stock · 2026-08-12