NVIDIA Groq 3 LPX hits 3,431 tokens/s output at 100K context in benchmarks
scaling01 · x · 2026-09-01
An NVIDIA technical blog covers Groq 3 LPX, the interactive inference accelerator for the Vera Rubin platform. Artificial Analysis benchmarked Gemma 4 31B at 100K context and measured a median of 3,431 output tokens/s; on the SPEED-Bench coding benchmark the system reached a median of 4,767 tokens/s with a P80 of 5,520.
The accelerator relies on a deterministic execution model with compiler-scheduled chip-to-chip communication and fine-grained compute-communication overlap. Groq 3 LPX pairs with Vera Rubin NVL72 in configurations such as prefill-decode disaggregation, attention-FFN disaggregation, and external-drafter speculative decoding.
More from Infra
- Spending $60k on Macs for Local LLMs Still Beats by $10 Cloud Subscription — leebase65 · 2026-09-01
- Call for agent infra: Who will build the open source Codex-style browser? — hwchase17 · 2026-09-01
- Qwen2.5-72B Local Benchmark: 35 vs 65 Tokens/s Configs Analyzed — LittleCelebration412 · 2026-09-01
- Why is NVIDIA MIG still limited to 7 instances on B300/Rubin? — StasBekman · 2026-09-01
- Polymarket: 69% Chance Any State Enacts Data Center Moratorium by 2026 — Polymarket · 2026-09-01
- NVLink adoption deepens investment in NVIDIA hardware — BenBajarin · 2026-09-01