NVIDIA Groq 3 LPX hits 3,431 tokens/s output at 100K context in benchmarks

scaling01 · x · 2026-09-01

An NVIDIA technical blog covers Groq 3 LPX, the interactive inference accelerator for the Vera Rubin platform. Artificial Analysis benchmarked Gemma 4 31B at 100K context and measured a median of 3,431 output tokens/s; on the SPEED-Bench coding benchmark the system reached a median of 4,767 tokens/s with a P80 of 5,520.

The accelerator relies on a deterministic execution model with compiler-scheduled chip-to-chip communication and fine-grained compute-communication overlap. Groq 3 LPX pairs with Vera Rubin NVL72 in configurations such as prefill-decode disaggregation, attention-FFN disaggregation, and external-drafter speculative decoding.

Original post →

More from Infra

Infra channel →