Nebius brings NVIDIA Groq 3 LPX online; Rubin NVL72 hits up to 30x throughput per MW on agentic coding

demian_ai · x · 2026-08-25

Nebius announced it is the first AI cloud to bring NVIDIA Groq 3 LPX online, putting ultra-fast inference on the same platform developers use to run leading open models in production.

The framing is the "AI factory": power is the scarce resource, and the metric is completed sequential agent steps inside a fixed megawatt — not peak FLOPS. Vera Rubin NVL72 reportedly delivers up to 30x higher throughput/MW on real agentic coding trajectories (context growth, tool calls, subagents) vs the previous generation.

The platform is codesigned for the full loop: Vera CPU feeds the GPUs with orchestration and tool use, Rubin does the heavy lifting, and LPX handles the low-latency generation step where sequential calls used to pile up.

Related event: NVIDIA's Groq 3 LPX Enters Full Production, Nebius First to Deploy(7 posts)→

Original post →

More from Infra

Infra channel →