NVIDIA Releases Groq 3 LPX for Agentic AI with 4x Faster Inference
NVIDIA Blog · rss · 2026-08-24
NVIDIA announced the production launch of NVIDIA Groq 3 LPX, an extension to the Vera Rubin NVL72 system designed for low-latency inference in the agentic era. In benchmarks running Gemma 4 31B, it delivered 3,400 tokens per second, 4x faster than the nearest alternative. SpaceXAI adopted NVIDIA Vera CPUs for agentic AI, while CoreWeave and Nebius deployed associated networking and acceleration tech.
More from Infra
- Intel China Follows Bonsai on Quantization; Quinary Quantization Bet to be on Pareto Frontier — georgejrjrjr · 2026-08-25
- Debunking "Data Centers Only Serve Billionaires" — AIandDesign · 2026-08-25
- Balance Enterprise Self-Hosting with Frontier Models for Optimal Strategy — TheZachMueller · 2026-08-25
- Mobile inference spans 30x: LFM fastest at 0.9s, Falcon slowest at 26.7s — ArtificialAnlys · 2026-08-25
- Mobile memory usage spans 19x: 9B models consume nearly 7GB peak — ArtificialAnlys · 2026-08-25
- Liquid AI Releases Pipette, a Benchmark Suite for On-Device AI — maximelabonne · 2026-08-25