Nvidia Claims Groq 3 LPX 4x Faster Than Cerebras, But Needs 64 Accelerators
The Decoder · rss · 2026-08-25
Nvidia is moving its Groq 3 LPX inference chip into full production, reporting 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras. However, The Register notes Nvidia needs at least 64 accelerators to achieve this, while Cerebras needs only one or two. How the architecture scales with large MoE models remains an open question.
More from Infra
- Apple Silicon Still Leads Single-Threaded Performance — lemire · 2026-08-25
- Unsloth AI aims for day-zero llama.cpp support for Qwen models — danielhanchen · 2026-08-25
- DevOps to AI Infra is becoming a serious career path — _jaydeepkarale · 2026-08-25
- AI Growth Forces Smarter Cloud: Power Becomes the New Bottleneck — DavidLinthicum · 2026-08-25
- Is Agent Collaboration the Next Major AI Infrastructure Layer? — Plenty-Ad-8268 · 2026-08-25
- Goldman: China's advanced chip self-sufficiency to reach 66% by 2035 — pstAsiatech · 2026-08-25