NVIDIA's Groq 3 LPX enters full production, hits 3,400 tokens/sec in tests

zephyr_z9 · x · 2026-08-24

NVIDIA says Groq 3 LPX, its low-latency inference accelerator designed to extend Vera Rubin NVL72, is now in full production.

In Artificial Analysis testing, Groq 3 LPX reached 3,400 output tokens/sec running Gemma 4 31B with a 100K-token context; NVIDIA claims 4x faster responsiveness than the nearest alternative for latency-sensitive agentic workloads.

The architecture splits inference between Rubin GPUs for large-scale context processing and LPX for fast token generation, targeting coding agents, multi-step reasoning, and tool-use workloads. Nebius will be the first AI cloud to deploy it via Nebius Token Factory.

Related event: NVIDIA Groq 3 LPX Enters Full Production with Record 3400 tokens/s Inference(2 posts)→

Original post →

More from Infra

Infra channel →