NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s
NVIDIA's Groq 3 LPX inference accelerator entered full production, delivering a record 3,400 tokens/s on Gemma 4 31B, with Nebius becoming the first cloud provider to deploy it under last year's Groq licensing deal.
2026-08-24 ~ 2026-08-25 · 4 related posts
- NVIDIA's Groq 3 LPX enters full production, hits 3,400 tokens/sec in tests — zephyr_z9 · 2026-08-24
- Nebius Adopts NVIDIA Groq 3 LPX for Fastest Inference — BenBajarin · 2026-08-25
- NVIDIA Confirms Groq LPX Production; Nebius First to Deploy in Cloud — IanAndrewsDC · 2026-08-25
1 near-duplicate retellings: nvidia