SketchSSM: 11x lower state traffic at rank 8 with near-lossless accuracy

sehoonkim418 · x · 2026-10-08

The SketchSSM authors shared additional numbers: at mean sketch rank 8, the method cuts linear-attention state traffic by 11x with near-lossless accuracy across reasoning and recall benchmarks, spanning Mamba-2, Gated DeltaNet, and KDA architectures.

The authors stress that pruning and quantization fall apart well before reaching comparable compression. With decode kernels up to 7.3x faster and 2.8x higher end-to-end throughput on Nemotron 3 Super, it ships as a drop-in vLLM install.

Related event: SketchSSM: Approximate State Reading Speeds Linear Attention Decoding by 2.8x(8 posts)→

Original post →

More from Infra

Infra channel →