SparseDecoding: Decoding-Aware Pruning Yields up to 1.48x Faster LLM Inference

encodelab · hf · 2026-10-09

SparseDecoding addresses a distribution shift in Hessian-guided pruning: calibration uses natural sequences while decoding consumes self-generated tokens. The framework builds calibration matrices from activations collected during dense autoregressive generation (excluding prefill) and ships an optimized N:M sparse matrix-vector kernel with bitmask indexing. On Llama-3.1-8B, Llama-3.3-70B, and Qwen3-14B/32B, it outperforms fixed-text calibration on long-form benchmarks with up to 1.48x end-to-end decoding speedup on A100.

Original post →

More from Infra

Infra channel →