NVIDIA Groq 3 LPX enters production with record 3,400 tokens/sec
nvidia · x · 2026-08-24
NVIDIA announced the full production of the Groq 3 LPX interactive AI inference accelerator. An extension of the Vera Rubin platform, it achieved a record 3,400 output tokens per second on the Gemma 4 31B benchmark, drastically accelerating agentic systems. Nebius is the first AI cloud to adopt the chip.
Related event: NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s(4 posts)→
More from Infra
- Can an M3 MacBook Pro with 18GB RAM Handle ComfyUI for AI Video? — Kevin_gato · 2026-08-25
- OneTriangleAI launches ultra-low latency DeepSeek V4 hosting — ycombinator · 2026-08-25
- Modular hits Pareto frontier for GLM-5.2 price-performance — clattner_llvm · 2026-08-25
- Nvidia executive claims data centers are now power limited at Hot Chips — firstadopter · 2026-08-25
- llama.cpp Resumes ROCm 7.14 Nightlies; Script Sought for 7900XTX — W61k3r · 2026-08-25
- Optimizing Context Engineering with Small Model Sub-Agents to Cut Costs — samsamy7 · 2026-08-25