Is Local LLM Prefill Throughput Severely Underestimated?
GabryIta · reddit · 2026-07-07
A developer points out that when discussing the ROI of running local LLMs, the focus is almost entirely on decoding (output) speed, while prefill (input) speed is ignored. Using the example of running GLM 5.2 (with 4-bit quantization, speculative decoding, and other optimizations) on 4 NVIDIA DGX Spark units, one can achieve about 60 output tokens/second with 6 concurrent requests. Theoretically, running 24/7 yields about 5.18 million output tokens a day, costing roughly $22/day at $4.40 per million output tokens. However, the prefill throughput for the same configuration is about 3000 tokens/second—50 times faster than output. Even with a lower prefill unit price (about $1.40 per million input tokens), this massive throughput gap should significantly impact ROI. The author questions why the industry rarely includes prefill in its evaluations.
More from Infra
- Intel 10-Q points to 18A/14A progress and “potential significant external customers” — BenBajarin · 2026-07-27
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27