NVIDIA Releases Groq 3 LPX for Agentic AI with 4x Faster Inference

NVIDIA Blog · rss · 2026-08-24

NVIDIA announced the production launch of NVIDIA Groq 3 LPX, an extension to the Vera Rubin NVL72 system designed for low-latency inference in the agentic era. In benchmarks running Gemma 4 31B, it delivered 3,400 tokens per second, 4x faster than the nearest alternative. SpaceXAI adopted NVIDIA Vera CPUs for agentic AI, while CoreWeave and Nebius deployed associated networking and acceleration tech.

Original post →

More from Infra

Infra channel →