NVIDIA Groq 3 LPX Hits 3,400+ tok/s Decoding on Small Models

Benchmarks show NVIDIA's Groq 3 LPX, which swaps HBM for 128GB of ultra-fast SRAM, achieves roughly 3,400 tokens/s decoding on the Gemma 4 31B model across 10k and 100k input sequences.

2026-08-27 ~ 2026-08-27 · 2 related posts