Tencent's HiLS Attention Boosts Long-Context Extrapolation 64x, Speeds Up Inference 15.7x
jiqizhixin · x · 2026-07-30
Tencent's HY Team introduced HiLS Attention (Hierarchical Sparse Attention), a new mechanism to tackle the massive GPU memory and compute demands of infinite-length contexts in LLMs.
The approach teaches models to focus only on key text chunks. It matches full attention performance on short tasks while extrapolating context 64 times further. Additionally, it delivers up to 15.7x faster inference, potentially marking a game changer for long-context AI modeling.
More from Models
- vLLM Announces Day 0 Support for Kimi K3 Across NVIDIA Architectures — vllm_project · 2026-07-30
- Leaked Opus 5 Benchmarks: Major Jumps in Research Math and Long-Context — echen · 2026-07-30
- Opus 5 benchmarks: #2 in research math and long-context agents, half the price — echen · 2026-07-30
- Kimi K3 Lands on DigitalOcean Powered by vLLM for Efficient Inference — vllm_project · 2026-07-30
- Moonshot Releases 2.8T-Parameter Kimi K3; Modal Achieves 460 TPS with Speculative Decoding — sarahcat21 · 2026-07-30
- Claude Is Down: Service Outage Confirmed by Status Page — gregsadetsky · 2026-07-30