ExLlamaV3 Is Underrated: Better Quants, Lower KLD, Faster Than llama.cpp
Embarrassed_Soup_279 · reddit · 2026-09-08
A Reddit user argues the local inference backend ExLlamaV3 (ExL3) is underrated: in personal tests it delivers higher-quality quants, lower KLD metrics, and faster speeds than llama.cpp (NVIDIA GPUs only). CPU MoE offload was only recently added, which may explain low adoption.
- Daily driver setup: tabbyAPI exl3 backend + qwen3 27B SC 6bpw H6 and qwen3 flash-next 4bpw
- The author stresses these are personal test sets, not formal benchmarks
- Encourages more people to try the project
More from Infra
- First third-party TPU inference benchmark: Google Ironwood up to 50% better perf/$ than B200 — dylan522p · 2026-09-08
- Qualcomm's Adreno Matrix Cores put AI acceleration inside the GPU pipeline for the first time — ryanshrout · 2026-09-08
- Polymarket puts 74% odds on a US state enacting a data center moratorium by end of 2026 — Polymarket · 2026-09-08
- Denver data center filmed heavily watering lawn while 1.5M residents face drought restrictions — Polymarket · 2026-09-08
- One mental model for Kubernetes, Slurm, Ray, and Spark: a unified take on distributed compute — ArchitectingAI · 2026-09-08
- Aurora Fork Fixes OpenCode API HTTP 400 Errors and Cuts Token Costs Up to 80% — entitybtw · 2026-09-08