vLLM Compressor v0.13.0 Introduces MoE Expert Pruning and Arbitrary Bit-Width Quantization
vllm_project · x · 2026-08-12
The vLLM project released LLM Compressor v0.13.0, highlighting major updates:
- REAP Expert Pruning: Structurally removes whole experts from MoE models based on calibration saliency scores, allowing subsequent FP8 or NVFP4 quantization of the remaining components.
- Arbitrary Bit-Width Quantization: Supports dense packing for non-power-of-2 bit widths (3, 5, 6, 7-bit) with no wasted bits, plus 16 new WxAy presets.
- Expanded Support: Broadens MoE linearization for more architectures (including Transformers v5.13.0 models) and improves Intel XPU compatibility.
More from Infra
- NVIDIA and Microsoft Optimize vLLM: 7.3x Faster Weights Loading on H100 — NVIDIAAI · 2026-08-12
- vLLM Teams Up with Microsoft and NVIDIA for Up to 7.3x Faster Model Loading on H100/A100 — vllm_project · 2026-08-12
- Skilled Labor Shortage Emerges as the Biggest Bottleneck for US AI Data Center Boom — coinfanking · 2026-08-12
- Bloom Energy Surges 14% as Data Centers Adopt On-Site Power for AI — BenBajarin · 2026-08-12
- Alibaba and Moonshot Target 10T Parameter Models, Hitting Memory Bottlenecks — pstAsiatech · 2026-08-12
- Base Power Raises $1B at $12B Valuation, Targeting AI Energy Dominance — Not Boring (Packy McCormick) · 2026-08-12