Nebius Adopts NVIDIA Groq 3 LPX for Fastest Inference

BenBajarin · x · 2026-08-25

Nebius AI Cloud has become the first to adopt the NVIDIA Groq 3 LPX, designed to accelerate token generation for NVIDIA Vera Rubin NVL72. Purpose-built for low-latency agentic workloads, the Groq 3 LPX rack features 256 LPU accelerators. Benchmarks show it running Gemma 4 31B at 3,400 tokens per second, the fastest performance ever recorded for that model.

Related event: NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s(4 posts)→

Original post →

More from Infra

Infra channel →