Tencent Hunyuan Open-Sources HiLS Sparse Attention Architecture

机器之心 · wechat · 2026-07-20

The Tencent Hunyuan team open-sourced HiLS-Attention, a novel hierarchical landmark sparse attention architecture. Through mathematical and structural innovations, it simultaneously resolves two major pain points of traditional sparse attention: insufficient expressiveness and non-differentiability end-to-end, breaking the dilemma of choosing between efficiency and performance.

Existing Pain Points: Traditional chunked sparse attention often uses mean or max pooling to estimate block importance, which introduces systematic mathematical bias. Conversely, parameterized summary methods use discrete Top-K selection operations that cut off gradient backpropagation, preventing end-to-end optimization of the selection process.

HiLS Architectural Innovations:

Experimental Results: Validated on 345M to 7B models, short-text performance is nearly on par with full attention. Trained with only an 8K context, it achieves zero-shot extrapolation to 4M lengths. At a 512K context, prefill and decode are accelerated by 13.5x and 15.7x respectively. Furthermore, thanks to compression denoising effects, the model even outperforms full attention mechanisms on certain long-text retrieval tasks.

Original post →

More from Models

Models channel →