Nvidia Claims Groq 3 LPX 4x Faster Than Cerebras, But Needs 64 Accelerators

The Decoder · rss · 2026-08-25

Nvidia is moving its Groq 3 LPX inference chip into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras. However, The Register notes Nvidia needs at least 64 accelerators to achieve this, while Cerebras needs only one or two. How the architecture scales with large MoE models remains an open question.

Original post →

More from Infra

Infra channel →