Tencent Hunyuan Open-Sources HiLS Sparse Attention Architecture
机器之心 · wechat · 2026-07-20
The Tencent Hunyuan team open-sourced HiLS-Attention, a novel hierarchical landmark sparse attention architecture. Through mathematical and structural innovations, it simultaneously resolves two major pain points of traditional sparse attention: insufficient expressiveness and non-differentiability end-to-end, breaking the dilemma of choosing between efficiency and performance. **Existing Pain Points**: Traditional chunked sparse attention often uses mean or max pooling to estimate block importance, which introduces systematic mathematical bias. Conversely, parameterized summary methods use discrete Top-K selection operations that cut off gradient backpropagation, preventing end-to-end optimization of the selection process. **HiLS Architectural Innovations**: - **Differentiable Scoring Function**: Utilizes first-order Taylor expansion to approximate block importance as "relevance term + entropy bias", adaptively unifying uniform and concentrated extreme distributions. - **Bridging Gradient Breakpoints**: Splits attention weights into intra-block and inter-block two-level Softmax, allowing proxy quality to participate directly in forward computation. This enables gradients from the language modeling Loss to directly guide the model in learning how to select blocks. **Experimental Results**: Validated on 345M to 7B models, short-text performance is nearly on par with full attention. Trained with only an 8K context, it achieves zero-shot extrapolation to 4M lengths. At a 512K context, prefill and decode are accelerated by 13.5x and 15.7x respectively. Furthermore, thanks to compression denoising effects, the model even outperforms full attention mechanisms on certain long-text retrieval tasks.
More from Models
- Claude 20x users report sharply tighter limits and faster quota burn — MarcJSchmidt · 2026-07-21
- Cola launches July, the latest model jokingly billed as “second only to Fable” — oran_ge · 2026-07-21
- Kimi K3 looks stronger and about 5× cheaper on a frontend dashboard task — OwariDa · 2026-07-21
- Last Week in AI recap: Anthropic’s $65B round, IPO filing, and Microsoft’s MAI push — Last Week in AI · 2026-07-21
- A user says Claude 4.6 felt worse yesterday and asks whether model quality can drift over time — Rahios · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21