SketchSSM cuts linear-attention state traffic 10x, speeds decode up to 7.3x on B300

sehoonkim418 · x · 2026-10-08

Researchers from Columbia and collaborators released SketchSSM, a new method targeting the recurrent-state-read bottleneck in hybrid-attention models: it keeps full-state updates but approximates state reads. The idea is to exploit low-rank state-weighted query approximation with offline-fixed basis vectors — at each state update, the full state is read once to precompute outputs stored in a compact sketch; subsequent decode steps reconstruct outputs from sketch vectors plus query-dependent coefficients.

Key results:

It's a drop-in for vLLM (pip install sketchssm), with paper, blog, code, and calibrations all released.

Related event: SketchSSM Cuts Linear Attention State Traffic 11x With Near-Lossless Accuracy(6 posts)→

Original post →

More from Infra

Infra channel →