SketchSSM cuts linear-attention decode latency up to 2.8x by reading a compact state sketch

sehoonkim418 · x · 2026-10-08

The author identifies a key bottleneck in hybrid LLMs: linear attention must read the full recurrent state from GPU memory at every decode step, accounting for up to 75% of decode latency at large batch sizes. Prior work ReplaySSM buffered updates and wrote state once per window, but reads remained — linear attention still took up to 53% of latency. SketchSSM addresses this by writing to the full state while reading from a compact sketch, achieving up to 2.8x faster decoding for state-space and linear-attention models without accumulating errors.

Related event: SketchSSM: Approximate State Reading Speeds Linear Attention Decoding by 2.8x(8 posts)→

Original post →

More from Infra

Infra channel →