Hybrid Linear Attention HLA Gains Up to 5.57 Points on LongBench-V2
Monash · hf · 2026-10-07
Monash researchers introduce Hybrid Linear Attention (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet addressing linear attention's difficulty in selectively accessing sparse, distant information after compressing history into recurrent states.
HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact self-attentively pooled representatives. Each gate interpolates the chunk's historical transition with the identity map, controlling both additive memory and transformation of earlier states, with effective-support regularization encouraging concentrated routing.
Across Qwen3.5 models from 0.8B to 9B, HLA beats native GDN and fixed chunk mixing by up to 5.57 points on LongBench-V2 and 3.97 on RULER; a from-scratch 1.3B model shows gains growing from 0.83 points at 4K to 4.22 at 32K context.
More from Research
- Slides released: 'Ultimate guide to multi-harness RL' talk from Kernel Panic Madrid — SergioPaniego · 2026-10-07
- AI Math Skeptics Pushed Back On—'Give It 2 More Months,' Researcher Quips — AvivTamar1 · 2026-10-07
- New proof bounds π's irrationality exponent at 6.0446, fully formalized in Lean 4 — Michael_D_Moor · 2026-10-07
- After 3-sum hits n^1.999, complexity theorists debate whether math is 'all square packing' — thomasahle · 2026-10-07
- NVIDIA's UNREAL: one model unifies corpus retrieval and long-context at 128K+ — nvidia · 2026-10-07
- gpt-4.1-nano as RAG judge barely beats chance: AUROC 0.603 on RAGBench — Giulio Zeloni · 2026-10-07