Nebius Adopts NVIDIA Groq 3 LPX for Fastest Inference
BenBajarin · x · 2026-08-25
Nebius AI Cloud has become the first to adopt the NVIDIA Groq 3 LPX, designed to accelerate token generation for NVIDIA Vera Rubin NVL72. Purpose-built for low-latency agentic workloads, the Groq 3 LPX rack features 256 LPU accelerators. Benchmarks show it running Gemma 4 31B at 3,400 tokens per second, the fastest performance ever recorded for that model.
Related event: NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s(4 posts)→
More from Infra
- Can an M3 MacBook Pro with 18GB RAM Handle ComfyUI for AI Video? — Kevin_gato · 2026-08-25
- OneTriangleAI launches ultra-low latency DeepSeek V4 hosting — ycombinator · 2026-08-25
- Modular hits Pareto frontier for GLM-5.2 price-performance — clattner_llvm · 2026-08-25
- Nvidia executive claims data centers are now power limited at Hot Chips — firstadopter · 2026-08-25
- llama.cpp Resumes ROCm 7.14 Nightlies; Script Sought for 7900XTX — W61k3r · 2026-08-25
- Optimizing Context Engineering with Small Model Sub-Agents to Cut Costs — samsamy7 · 2026-08-25