NVIDIA Groq 3 LPX Hits 3,431 tok/s on Gemma 4 Long Context Benchmark
ricklamers · x · 2026-08-25
NVIDIA released a technical blog detailing the performance of the Groq 3 LPX accelerator on the Vera Rubin platform. Tested with the Gemma 4 31B model, the system achieved a median output speed of 3,431 tokens/second over a 100K context window, peaking at 10,996 tokens/second in high AL cases. The architecture leverages deterministic compiler scheduling, fine-grained compute-communication overlap, and pre-planned networking to minimize first-bit latency. Additionally, Groq 3 LPX supports various co-execution configurations—such as prefill-decode disaggregation and speculative external-drafter decoding—aiming to deliver best-in-class energy efficiency and interactivity for trillion-parameter models.
Related event: NVIDIA Showcases Gemma 4 Inference Breaking 10K Speed on LPX(2 posts)→
More from Infra
- Analyst expects TPU shipments to surpass NVIDIA's by 2028 — AccBalanced · 2026-08-25
- Guide: Running Hermes Agent on a Raspberry Pi — LeviTurk · 2026-08-25
- West Virginia targets data centers; proximity to nuclear reactors cited as a key advantage — mimi10v3 · 2026-08-25
- JetBrains Local AI Uses Qwen3.6 27B for Optimization — Danmoreng · 2026-08-25
- SpaceX plans million-satellite constellation with Nvidia Vera Rubin compute, scaling Grok to 10GW — ns123abc · 2026-08-25
- Weaviate Adds Configurable Effort Parameter to Scale Test-Time Compute in Search Mode — CShorten30 · 2026-08-25