NVIDIA Groq 3 LPX Hits 3,431 tok/s on Gemma 4 Long Context Benchmark

ricklamers · x · 2026-08-25

NVIDIA released a technical blog detailing the performance of the Groq 3 LPX accelerator on the Vera Rubin platform. Tested with the Gemma 4 31B model, the system achieved a median output speed of 3,431 tokens/second over a 100K context window, peaking at 10,996 tokens/second in high AL cases. The architecture leverages deterministic compiler scheduling, fine-grained compute-communication overlap, and pre-planned networking to minimize first-bit latency. Additionally, Groq 3 LPX supports various co-execution configurations—such as prefill-decode disaggregation and speculative external-drafter decoding—aiming to deliver best-in-class energy efficiency and interactivity for trillion-parameter models.

Related event: NVIDIA Showcases Gemma 4 Inference Breaking 10K Speed on LPX(2 posts)→

Original post →

More from Infra

Infra channel →