NVIDIA's Groq 3 LPX enters full production, hits 3,400 tokens/sec in tests
zephyr_z9 · x · 2026-08-24
NVIDIA says Groq 3 LPX, its low-latency inference accelerator designed to extend Vera Rubin NVL72, is now in full production.
In Artificial Analysis testing, Groq 3 LPX reached 3,400 output tokens/sec running Gemma 4 31B with a 100K-token context; NVIDIA claims 4x faster responsiveness than the nearest alternative for latency-sensitive agentic workloads.
The architecture splits inference between Rubin GPUs for large-scale context processing and LPX for fast token generation, targeting coding agents, multi-step reasoning, and tool-use workloads. Nebius will be the first AI cloud to deploy it via Nebius Token Factory.
More from Infra
- Maxime Labonne shares study notes on speculative decoding — maximelabonne · 2026-08-25
- Enterprise AI cost control: blending self-hosted open models with frontier APIs — TheZachMueller · 2026-08-25
- SpaceXAI Adopts NVIDIA Vera CPU for Agentic AI and Space Deployment — nvidia · 2026-08-25
- Protocol Labs cuts funding for Shipyard, IPFS operations to wind down — jedisct1 · 2026-08-25
- Stripe to Acquire OpenRouter as Token Usage Hits 4.5+ Quadrillion Annual Rate — rohanpaul_ai · 2026-08-25
- OpenRouter usage surges 9,000x; agent workloads dominate token consumption — rohanpaul_ai · 2026-08-25