SketchSSM cuts SSM state traffic 11x at sketch rank 8 with near-lossless accuracy
sehoonkim418 · x · 2026-10-08
- Headline result: At mean sketch rank 8, SketchSSM reduces state traffic in SSM-style models by 11x while keeping near-lossless accuracy on multiple reasoning and recall benchmarks. The author notes pruning and quantization fall apart well before that point.
- Generality: The method works across Mamba-2, Gated DeltaNet, and KDA architectures.
- How it works: Within a window the state is fixed and only queries change, so queries are compressed onto a small low-rank basis and the state's outputs are computed once for that basis. Every query is a mix of the basis, so its output is just the same mix of the stored outputs—no full-state read required.
More from Infra
- Chrome's new DecisionModel API reverse-engineered: prompts, limits and engine tests — dejanseo · 2026-10-08
- China's Power Glut Meets Data Centers; Immersion Cooling Traced to Bitcoin Miners — teortaxesTex · 2026-10-08
- omarchy-cluster runs the full 753B-param GLM-5.3 across four old Macs as one endpoint — natesiggard · 2026-10-08
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08