SketchSSM cuts linear-attention state traffic 10x, speeds decode up to 7.3x on B300
sehoonkim418 · x · 2026-10-08
Researchers from Columbia and collaborators released SketchSSM, a new method targeting the recurrent-state-read bottleneck in hybrid-attention models: it keeps full-state updates but approximates state reads. The idea is to exploit low-rank state-weighted query approximation with offline-fixed basis vectors — at each state update, the full state is read once to precompute outputs stored in a compact sketch; subsequent decode steps reconstruct outputs from sketch vectors plus query-dependent coefficients.
Key results:
- Across Mamba-2, GDN, and KDA models, at mean sketch rank 8 it cuts state-access traffic 10x while matching FP32 full-state accuracy on four decode benchmarks and preserving recall on four RULER retrieval tasks.
- On a single NVIDIA B300, linear-attention kernels run up to 7.30x faster than standard vLLM, with 2.8x higher end-to-end decode throughput on Nemotron 3 Super.
- The authors note pruning and quantization degrade well before reaching comparable compression.
It's a drop-in for vLLM (pip install sketchssm), with paper, blog, code, and calibrations all released.
More from Infra
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08
- DWDM wavelength lock tightens from ±12.5GHz to ±3.5GHz, pushing optical interconnect costs — jwt0625 · 2026-10-08
- llama.cpp merges Metal kernel PR covering all 26 weight formats, up to 4.4x faster MMA on Mac — ggerganov · 2026-10-08
- Theo: viral $132M/year token cost claim is wrong — closer to $3M now, $1.2k soon — dotey · 2026-10-08