Gemma 4 31B hits 3,431 tokens/s on NVIDIA Groq 3
GlennCameronjr · x · 2026-08-27
Benchmarks on NVIDIA's Groq 3 LPX accelerator show Gemma 4 31B achieving a median output of 3,400 tokens/s across 10k and 100k input sequences, setting a new speed record for the model at 100k context.
Related event: NVIDIA Groq 3 LPX Hits 3,400+ tok/s Decoding on Small Models(2 posts)→
More from Infra
- NVIDIA FLARE Cuts Federated VLM Training Traffic by 99% — dl_weekly · 2026-08-27
- Open Source AI Share Hits 62% on Vercel, Eclipsing Closed Source Models — gajesh · 2026-08-27
- Same Budget: 256GB Mac or Two DGX Sparks for 70B Inference? — Whyme-__- · 2026-08-27
- Edviro Builds World Model to Unify Data Center Operations — ycombinator · 2026-08-27
- Chinese Models Top US in Token Usage on OpenRouter; Efficiency Becomes Advantage — AccBalanced · 2026-08-27
- SandboxAQ Open-Sources Switch for Shared AI-Agent Workspaces — Codeblix_Ltd · 2026-08-27