LLM Compressor v0.14.0 makes GPTQ quantization up to 30x faster with new Triton kernel
vllm_project · x · 2026-09-24
The vLLM project released LLM Compressor v0.14.0, delivering the biggest GPTQ speedup since launch.
Key updates:
- A new Triton-based quantization kernel makes GPTQ roughly 15x faster end-to-end; batching layers that share a shape pushes this to 30x on some MoE workloads, and the remaining eager path got an independent 1.5-2x boost. Hessian offloading was removed.
- Expanded MSE/iMatrix observers widen the grid search space (a superset of the Fourosix-style strategy) and beat GPTQ for NVFP4 on internal benchmarks; a new Triton kernel makes MSE observation 10x faster with bitwise parity.
- REAP pruning gains distributed DDP support plus an e-Score correction.
- Support added for GLM 5.3 and Qwen3.8.
Release notes are on GitHub (vllm-project/llm-compressor, 3.8k stars).
More from Infra
- Google's Project Suncatcher to launch prototype datacenter satellite on Falcon 9 Oct 1 — McDonaghMatthew · 2026-09-25
- Lightmatter CEO: moving lasers onto 300mm silicon wafers to unlock optical interconnect — BenBajarin · 2026-09-25
- Oracle Says Datacenter Force-Majeure Notice Doesn't Mean Project Delay — ns123abc · 2026-09-25
- 100MW+ Datacenters Wait 5-10 Years for Grid Power; Cato Says Let Firms Build Their Own — McDonaghMatthew · 2026-09-25
- Four scaling paradigms keep postponing the wall: params, CoT, recurrent depth, multi-agent — gordic_aleksa · 2026-09-25
- BofA models Meta deploying 5-6GW owned capacity in 2027 at $200B cost — Beth_Kindig · 2026-09-25