SketchSSM: 11x lower state traffic at rank 8 with near-lossless accuracy
sehoonkim418 · x · 2026-10-08
The SketchSSM authors shared additional numbers: at mean sketch rank 8, the method cuts linear-attention state traffic by 11x with near-lossless accuracy across reasoning and recall benchmarks, spanning Mamba-2, Gated DeltaNet, and KDA architectures.
The authors stress that pruning and quantization fall apart well before reaching comparable compression. With decode kernels up to 7.3x faster and 2.8x higher end-to-end throughput on Nemotron 3 Super, it ships as a drop-in vLLM install.
More from Infra
- Liquid AI ships open d1 decision models; 3B on RTX 3090 beats OpenAI API 8ms vs 345ms — philipvollet · 2026-10-08
- China's Power Glut Meets Data Centers; Immersion Cooling Traced to Bitcoin Miners — teortaxesTex · 2026-10-08
- omarchy-cluster runs the full 753B-param GLM-5.3 across four old Macs as one endpoint — natesiggard · 2026-10-08
- Only Samsung HBM meets Nvidia Vera Rubin performance requirements, per leak — zephyr_z9 · 2026-10-08
- Box CEO on agent compute: one app serving 100M users would need $2.8B in infra — inductionheads · 2026-10-08
- Firmus, valued near $44bn, may shelve ASX IPO as investors balk — nordicinst · 2026-10-08