Tencent releases Simple-Attention-Sparsification: sparse attention for long-context Qwen3-14B
tencent · hf · 2026-09-14
Tencent released Simple-Attention-Sparsification (SAS) on Hugging Face, a text-generation model finetuned from Qwen3-14B.
- Focus: sparse attention for long-context inference, built on an attention sparsification method
- Trained on the open-r1/OpenR1-Math-220k math reasoning dataset
- Ships with a paper (arxiv:2609.13141) and sglang inference support
Related event: Tencent Open-Sources SAPS Sparse Attention Method for Qwen3(2 posts)→
More from Models
- New Fable 5.1 Ultracode and Astra 6 Ultra models now running in parallel — scaling01 · 2026-09-15
- Hyperstition claims 62% pretraining cost cut and 1.7B math model beating Qwen3 — nick_linck · 2026-09-15
- Anthropic: Reward Hacking in Production RL Can Cause Natural Emergent Misalignment — eigenron · 2026-09-15
- Yandex releases AliceAI-T5-35B-A0.6B, a 35B MoE model activating just 0.6B params — yandex · 2026-09-15
- OpenAI Agents Caught Using 10+ Unauthorized Websites as Secret Communication Channels — Originalboy69 · 2026-09-15
- Why an open model seizing a compute stockpile is implausible: defenders get the best models too — teortaxesTex · 2026-09-15