SketchSSM: Approximate State Reading Speeds Linear Attention Decoding by 2.8x
A team from Columbia University and other institutions has released SketchSSM, tackling the memory-read bottleneck of linear attention decoding in hybrid-architecture LLMs: with an average sketch rank of 8, it cuts state-access traffic by 11x and speeds up decoding by 2.8x (other figures report up to 7.3x), with near-lossless accuracy across multiple reasoning and memory benchmarks, covering models such as Mamba-2 and Gated DeltaNet. Since reading state is the dominant cost in large-batch decoding, the method has direct practical value for long-horizon reasoning and high-concurrency serving.
Confirmed
- Background: in hybrid LLMs, linear attention must read the full recurrent state from GPU memory at every decoding step, which can account for up to 75% of decoding latency at large batch sizes; the earlier ReplaySSM mitigated write overhead by buffering updates and writing state once per multi-step window, but reads remained, with linear attention still accounting for about 53% of latency.
- Core idea: keep full state updates and only approximate state reads. It uses a low-rank state-weighted query approximation, with basis vectors fixed offline and one full-state read per state update to precompute outputs.
- Mechanism details (per the authors): within a window the state is unchanged and only the queries vary, so queries are compressed onto a small low-rank basis; the state output is computed once per window for that basis, and subsequent queries are combinations of the basis whose outputs are the same combination of stored outputs — no full-state read needed within the window.
- Error control: each window updates the full state once and refreshes a compact sketch within the same kernel; within the window, each step reads only the sketch, and compression error is never written back into the state. The authors note that pruning/quantization, which directly compress the state, let errors accumulate step by step, and long-horizon reasoning tasks show clear accuracy drops even with 2–4x traffic reduction.
- Results: at rank 8, state-read traffic drops 11x with near-lossless accuracy on multiple reasoning and retrieval benchmarks; decoding is 2.8x faster, with the paper also reporting up to 7.3x speedup.
Why it matters
- The linear-attention read bottleneck is a major latency source in large-batch serving; SketchSSM offers an acceleration path that neither sacrifices state exactness nor accumulates error, directly relevant for deploying Mamba-style and hybrid-attention models on long-horizon reasoning.
2026-10-08 ~ 2026-10-08 · 8 related posts
Primary sources
- SketchSSM cuts linear-attention state traffic 10x, speeds decode up to 7.3x on B300 — sehoonkim418 ·
- SketchSSM: 11x lower state traffic at rank 8 with near-lossless accuracy — sehoonkim418 ·
- SketchSSM speeds up Mamba and linear-attention decoding up to 2.8x by reading a compact sketch — sehoonkim418 ·
- [source] SketchSSM speeds up Mamba and linear-attention decoding up to 2.8x by reading a compact sketch — sehoonkim418 · 2026-10-08
- SketchSSM cuts linear-attention decode latency up to 2.8x by reading a compact state sketch — sehoonkim418 · 2026-10-08
- Linear attention takes up to 75% of decode latency at large batch, authors say — sehoonkim418 · 2026-10-08
- SketchSSM: approximate reads not the state, avoiding compounding compression errors — sehoonkim418 · 2026-10-08
- SketchSSM explained: low-rank query basis avoids full recurrent-state reads per decode step — sehoonkim418 · 2026-10-08
- SketchSSM cuts SSM state traffic 11x at sketch rank 8 with near-lossless accuracy — sehoonkim418 · 2026-10-08
- [source] SketchSSM: 11x lower state traffic at rank 8 with near-lossless accuracy — sehoonkim418 · 2026-10-08
- [source] SketchSSM cuts linear-attention state traffic 10x, speeds decode up to 7.3x on B300 — sehoonkim418 · 2026-10-08