HiLS-Attention 7B Released
pmttyji · reddit · 2026-07-10
**Tencent/HiLS-Attention-7B** has been released on Hugging Face, alongside the paper "**Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling**." This is a continued training checkpoint based on **OLMo3-7B**, focusing on efficient long-context modeling. The core ideas proposed in the paper are: - Traditional block sparse attention requires computing full chunk mass upfront, which is computationally expensive - HiLS-Attention uses **compressed chunk keys** to estimate chunk quality - It then decomposes attention into a two-level softmax: **inter-chunk** and **intra-chunk** - This allows sparse attention to be learned end-to-end under **next-token prediction loss** The repo also notes that this is a **pretrained base model** without alignment or safety fine-tuning, meaning it may contain biases and produce unsafe content.
Related event: Tencent Releases HiLS-Attention 7B Model(2 posts)→
More from Infra
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21
- More open models and llama.cpp updates are coming, says Merve Noyan — mervenoyann · 2026-07-21
- Why adding a second LLM provider breaks more than the API surface — Ok_Extension6373 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21