exllamav3 CPU-offload beats llama.cpp 3.2x prefill, 2x decode on Qwen
Lowkey_LokiSN · reddit · 2026-09-08
- On a 2×RTX 3080 20GB + 128GB DDR4 + Xeon 6148 setup, llama.cpp running Unsloth's Q4KXL quant of Qwen-3.8-Flash-Next managed 270tps prefill and 13tps decode (degrading with context).
- exllamav3's 4.05 EXL3 quant hit 25tps average decode (peaks of 32tps) stable through 160k context, with 870tps prefill — 3.2x faster prefill, 2x faster decode, and better quant quality.
- The win doesn't generalize: GLM 5.3 Flash's EXL3 ran 2x slower than llama.cpp, with AVX2 as the bottleneck; results depend on model architecture and hardware.
- Practical notes: exllamav3 + TabbyAPI is trickier to configure, and decode speeds warm up over a few thousand tokens (12tps rising to 25-30) as hot/cold experts calibrate.
More from Infra
- Polymarket puts 74% odds on a US state enacting a data center moratorium by end of 2026 — Polymarket · 2026-09-08
- NVIDIA reportedly closed $12.93B all-cash acquisition of Hugging Face — ryanshrout · 2026-09-08
- Denver data center filmed heavily watering lawn while 1.5M residents face drought restrictions — Polymarket · 2026-09-08
- One mental model for Kubernetes, Slurm, Ray, and Spark: a unified take on distributed compute — ArchitectingAI · 2026-09-08
- Aurora Fork Fixes OpenCode API HTTP 400 Errors and Cuts Token Costs Up to 80% — entitybtw · 2026-09-08
- NVIDIA Launches Free Tool to Turn Your PC Into a Personal AI Data Center — ohiocodernumerouno · 2026-09-08