SketchSSM: Approximate State Reading Speeds Linear Attention Decoding by 2.8x

A team from Columbia University and other institutions has released SketchSSM, tackling the memory-read bottleneck of linear attention decoding in hybrid-architecture LLMs: with an average sketch rank of 8, it cuts state-access traffic by 11x and speeds up decoding by 2.8x (other figures report up to 7.3x), with near-lossless accuracy across multiple reasoning and memory benchmarks, covering models such as Mamba-2 and Gated DeltaNet. Since reading state is the dominant cost in large-batch decoding, the method has direct practical value for long-horizon reasoning and high-concurrency serving.

Confirmed

Why it matters

2026-10-08 ~ 2026-10-08 · 8 related posts

Primary sources