ALHR sparse attention reads 30 keys per query, achieves 35x KV compression

Alarming-Emotion-894 · reddit · 2026-10-09

A developer released ALHR (Adaptive Learnable Hierarchical Routing), a tree-based sparse attention system achieving sub-quadratic inference while retaining accuracy.

Method

MQAR test at 1024 tokens

Limitations: training remains quadratic, inference is N log N; full-scale tests not yet done. Code and logs are open source on GitHub.

Related event: ALHR: Binary-Tree Sparse Attention Achieves O(NlogN) Long-Context Inference(2 posts)→

Original post →

More from Research

Research channel →