NVIDIA Groq 3 LPX Hits 3,400+ tok/s Decoding on Small Models
Benchmarks show NVIDIA's Groq 3 LPX, which swaps HBM for 128GB of ultra-fast SRAM, achieves roughly 3,400 tokens/s decoding on the Gemma 4 31B model across 10k and 100k input sequences.
2026-08-27 ~ 2026-08-27 · 2 related posts
- Gemma 4 31B hits 3,431 tokens/s on NVIDIA Groq 3 — GlennCameronjr · 2026-08-27
- Nvidia Groq 3 LPX system hits >3,400 tok/s on small models — Jsevillamol · 2026-08-27