NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s

NVIDIA's Groq 3 LPX inference accelerator entered full production, delivering a record 3,400 tokens/s on Gemma 4 31B, with Nebius becoming the first cloud provider to deploy it under last year's Groq licensing deal.

2026-08-24 ~ 2026-08-25 · 4 related posts

1 near-duplicate retellings: nvidia