LLM Compressor v0.14.0 makes GPTQ quantization up to 30x faster with new Triton kernel

vllm_project · x · 2026-09-24

The vLLM project released LLM Compressor v0.14.0, delivering the biggest GPTQ speedup since launch.

Key updates:

Release notes are on GitHub (vllm-project/llm-compressor, 3.8k stars).

Original post →

More from Infra

Infra channel →