Nvidia Groq 3 LPX system hits >3,400 tok/s on small models
Jsevillamol · x · 2026-08-27
Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s) using 128GB of ultrafast SRAM instead of HBM. Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models.
Related event: NVIDIA Groq 3 LPX Hits 3,400+ tok/s Decoding on Small Models(2 posts)→
More from Infra
- Economics of Becoming an OpenRouter Provider with H200 Nodes — ell-hol1 · 2026-08-27
- Analysis layers exist between Nvidia sales and enterprise ROI — iamKierraD · 2026-08-27
- Apple still leads in laptop processor performance years after M1, leaving Intel and AMD behind — lemire · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27
- Could Iceland Become Europe's AI Powerhouse? — SM_stories · 2026-08-27
- Reducing Agent Token Usage by 84% via Command Output Filtering — hheadshott · 2026-08-27