NVIDIA Showcases Gemma 4 Inference Breaking 10K Speed on LPX
NVIDIA's LPX team published benchmarks showing Gemma 4 exceeding 10K OTSU in high-quality settings, and Groq 3 LPX reaching 3431 tok/s with Gemma 4 31B on a 100K context on the Vera Rubin platform.
2026-08-25 ~ 2026-08-25 · 2 related posts
- NVIDIA blog: Gemma 4 hits 10,996 OTSU with Vera Rubin optimizations — ricklamers · 2026-08-25
- NVIDIA Groq 3 LPX Hits 3,431 tok/s on Gemma 4 Long Context Benchmark — ricklamers · 2026-08-25