Tencent releases SAPS: sparse attention that learns KV block selection end-to-end

teortaxesTex · x · 2026-09-15

Tencent has released a new sparse attention method, Simple Attention Sparsification (SAPS), on Hugging Face.

The key idea: instead of relying on handcrafted heuristics for which KV blocks to attend to, SAPS lets the language modeling loss directly optimize block selection per query — making sparse attention end-to-end learnable and potentially cutting long-context inference cost.

Related event: Tencent Open-Sources SAPS Sparse Attention Method for Qwen3(2 posts)→

Original post →

More from Research

Research channel →