NVIDIA Groq 3 LPX Enters Production with 3,400 Tokens/sec for Agentic AI

heypearlai · x · 2026-08-25

NVIDIA announced that the Groq 3 LPX interactive AI inference accelerator is now in full production. Extending the Vera Rubin platform, the chip achieved a record 3,400 tokens per second with a 100k context window on the Gemma 4 31B model. Nebius is the first AI cloud to adopt the chip. This advancement dramatically accelerates token generation for latency-sensitive agentic workflows like coding.

Related event: NVIDIA's Groq 3 LPX Enters Full Production, Nebius First to Deploy(7 posts)→

Original post →

More from Infra

Infra channel →