SketchSSM speeds up Mamba and linear-attention decoding up to 2.8x by reading a compact sketch
sehoonkim418 · x · 2026-10-08
- Problem: State-space and linear-attention models maintain a fixed-size recurrent state, but reading the full state at every decode step is a major bottleneck; pruning/quantization compounds errors and even 2–4× traffic reduction hurts long-reasoning accuracy.
- Method: SketchSSM keeps full-state updates while approximating reads. Low-rank state-weighted query approximation preserves read outputs, and the basis vectors can be fixed offline. At each state update, the full state is read once to precompute outputs for these basis vectors into a compact sketch; each subsequent step reconstructs outputs from query-dependent coefficients without a full-state read.
- Results: Across Mamba-2, GDN, and KDA models, a mean sketch rank of 8 cuts state-access traffic 10× while matching FP32 full-state accuracy on four decode benchmarks and preserving RULER recall. On one NVIDIA B300, linear-attention kernel speedups over standard vLLM reach 7.30× and 5.02×, with up to 2.8× overall decoding speedup.
- Paper: arXiv 2609.33051 (Omin Kwon et al., with Kurt Keutzer).
More from Infra
- Chrome's new DecisionModel API reverse-engineered: prompts, limits and engine tests — dejanseo · 2026-10-08
- China's Power Glut Meets Data Centers; Immersion Cooling Traced to Bitcoin Miners — teortaxesTex · 2026-10-08
- omarchy-cluster runs the full 753B-param GLM-5.3 across four old Macs as one endpoint — natesiggard · 2026-10-08
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08