ALHR sparse attention reads 30 keys per query, achieves 35x KV compression
Alarming-Emotion-894 · reddit · 2026-10-09
A developer released ALHR (Adaptive Learnable Hierarchical Routing), a tree-based sparse attention system achieving sub-quadratic inference while retaining accuracy.
Method
- Static binary trees plus learnable routing functions minimize keys read per query; a dense teacher is used in training phase 1.
MQAR test at 1024 tokens
- Keys read per query: dense 512 vs ALHR 30.
- Top-1 accuracy: 94.9% dense vs 92.1% ALHR.
- KV compression: 35.3x (2.83% read); 100% cache compression.
- Peak VRAM: dense 57MB (quadratic) vs ALHR 422MB (linear).
Limitations: training remains quadratic, inference is N log N; full-scale tests not yet done. Code and logs are open source on GitHub.
Related event: ALHR: Binary-Tree Sparse Attention Achieves O(NlogN) Long-Context Inference(2 posts)→
More from Research
- Google and Tel Aviv researchers unveil SepGen, generating video with per-source stems for 4D spatial audio — YonatanBitton · 2026-10-11
- Why OpenAI's quasi-Riemann result must cover all L-functions, vindicating Hardy's 1921 conjecture — burny_tech · 2026-10-11
- Schmidhuber Points to Section 20 of His 'Annotated History of Modern AI and Deep Learning' — SchmidhuberAI · 2026-10-11
- Stephen Wolfram's 'A New Kind of Science' Is Freely Readable Online — burny_tech · 2026-10-11
- Task-structured modularity emerges in AI networks, aligning with brain architecture — lulzxdxdxd · 2026-10-11
- Toward an Automated Science of the Mind: AI Enters Every Stage of Cognitive Research — burny_tech · 2026-10-11