Tencent's HiLS Attention Boosts Long-Context Extrapolation 64x, Speeds Up Inference 15.7x

jiqizhixin · x · 2026-07-30

Tencent's HY Team introduced HiLS Attention (Hierarchical Sparse Attention), a new mechanism to tackle the massive GPU memory and compute demands of infinite-length contexts in LLMs.

The approach teaches models to focus only on key text chunks. It matches full attention performance on short tasks while extrapolating context 64 times further. Additionally, it delivers up to 15.7x faster inference, potentially marking a game changer for long-context AI modeling.

Original post →

More from Models

Models channel →