NVIDIA Groq 3 LPX enters production with record 3,400 tokens/sec

nvidia · x · 2026-08-24

NVIDIA announced the full production of the Groq 3 LPX interactive AI inference accelerator. An extension of the Vera Rubin platform, it achieved a record 3,400 output tokens per second on the Gemma 4 31B benchmark, drastically accelerating agentic systems. Nebius is the first AI cloud to adopt the chip.

Related event: NVIDIA's Groq 3 LPX Enters Full Production at 3,400 Tokens/s(4 posts)→

Original post →

More from Infra

Infra channel →