SketchSSM cuts linear-attention decode latency up to 2.8x by reading a compact state sketch
sehoonkim418 · x · 2026-10-08
The author identifies a key bottleneck in hybrid LLMs: linear attention must read the full recurrent state from GPU memory at every decode step, accounting for up to 75% of decode latency at large batch sizes. Prior work ReplaySSM buffered updates and wrote state once per window, but reads remained — linear attention still took up to 53% of latency. SketchSSM addresses this by writing to the full state while reading from a compact sketch, achieving up to 2.8x faster decoding for state-space and linear-attention models without accumulating errors.
More from Infra
- Chrome's new DecisionModel API reverse-engineered: prompts, limits and engine tests — dejanseo · 2026-10-08
- China's Power Glut Meets Data Centers; Immersion Cooling Traced to Bitcoin Miners — teortaxesTex · 2026-10-08
- omarchy-cluster runs the full 753B-param GLM-5.3 across four old Macs as one endpoint — natesiggard · 2026-10-08
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08