ALHR: Binary-Tree Sparse Attention Achieves O(NlogN) Long-Context Inference
A developer introduced ALHR, a binary-tree-based sparse attention system that routes keys with learnable functions so each query reads only 30 keys, achieving O(NlogN) inference with 35x KV compression while retaining about 97% accuracy.
2026-10-09 ~ 2026-10-11 · 2 related posts
- ALHR sparse attention reads 30 keys per query, achieves 35x KV compression — Alarming-Emotion-894 · 2026-10-09
- O(NlogN) tree-based attention retains 97% accuracy on long-context MQAR benchmark — Alarming-Emotion-894 · 2026-10-11