Is Local LLM Prefill Throughput Severely Underestimated?
GabryIta · reddit · 2026-07-07
A developer points out that when discussing the ROI of running local LLMs, the focus is almost entirely on decoding (output) speed, while prefill (input) speed is ignored. Using the example of running GLM 5.2 (with 4-bit quantization, speculative decoding, and other optimizations) on 4 NVIDIA DGX Spark units, one can achieve about 60 output tokens/second with 6 concurrent requests. Theoretically, running 24/7 yields about 5.18 million output tokens a day, costing roughly $22/day at $4.40 per million output tokens. However, the prefill throughput for the same configuration is about 3000 tokens/second—50 times faster than output. Even with a lower prefill unit price (about $1.40 per million input tokens), this massive throughput gap should significantly impact ROI. The author questions why the industry rarely includes prefill in its evaluations.
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11