Nebius brings NVIDIA Groq 3 LPX online; Rubin NVL72 hits up to 30x throughput per MW on agentic coding
demian_ai · x · 2026-08-25
Nebius announced it is the first AI cloud to bring NVIDIA Groq 3 LPX online, putting ultra-fast inference on the same platform developers use to run leading open models in production.
The framing is the "AI factory": power is the scarce resource, and the metric is completed sequential agent steps inside a fixed megawatt — not peak FLOPS. Vera Rubin NVL72 reportedly delivers up to 30x higher throughput/MW on real agentic coding trajectories (context growth, tool calls, subagents) vs the previous generation.
The platform is codesigned for the full loop: Vera CPU feeds the GPUs with orchestration and tool use, Rubin does the heavy lifting, and LPX handles the low-latency generation step where sequential calls used to pile up.
Related event: NVIDIA's Groq 3 LPX Enters Full Production, Nebius First to Deploy(7 posts)→
More from Infra
- West Virginia targets data centers; proximity to nuclear reactors cited as a key advantage — mimi10v3 · 2026-08-25
- Nvidia calls Agentic AI the most complex computing workload in history — AccBalanced · 2026-08-25
- JetBrains Local AI Uses Qwen3.6 27B for Optimization — Danmoreng · 2026-08-25
- SpaceX plans million-satellite constellation with Nvidia Vera Rubin compute, scaling Grok to 10GW — ns123abc · 2026-08-25
- Weaviate Adds Configurable Effort Parameter to Scale Test-Time Compute in Search Mode — CShorten30 · 2026-08-25
- Qualcomm Acquires Modular to Build Open Software Stack for Heterogeneous Compute — clattner_llvm · 2026-08-25