SparseDecoding: Decoding-Aware Pruning Yields up to 1.48x Faster LLM Inference
encodelab · hf · 2026-10-09
SparseDecoding addresses a distribution shift in Hessian-guided pruning: calibration uses natural sequences while decoding consumes self-generated tokens. The framework builds calibration matrices from activations collected during dense autoregressive generation (excluding prefill) and ships an optimized N:M sparse matrix-vector kernel with bitmask indexing. On Llama-3.1-8B, Llama-3.3-70B, and Qwen3-14B/32B, it outperforms fixed-text calibration on long-form benchmarks with up to 1.48x end-to-end decoding speedup on A100.
More from Infra
- OpenAI revenue definition gap sparks AI hardware selloff as TSMC posts +55% YoY — tengyanAI · 2026-10-09
- LoRA over GGUF: fine-tune Qwen3.8-Flash-Next in 40 GiB VRAM without CPU offloading — woct0rdho · 2026-10-09
- Nunchux inference engine joins AMD's AI Inference Engines & Services ecosystem — junyanz89 · 2026-10-09
- tinygrad runs its GitHub Actions CI on 4 new tinybox machines — AIFlow_ML · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- TokenRouter Serving System Boosts Token-Level LLM Routing Throughput up to 64x — nics-efc · 2026-10-09