Hybrid Linear Attention HLA Gains Up to 5.57 Points on LongBench-V2

Monash · hf · 2026-10-07

Monash researchers introduce Hybrid Linear Attention (HLA), a query-dependent chunk-level attention mechanism for Gated DeltaNet addressing linear attention's difficulty in selectively accessing sparse, distant information after compressing history into recurrent states.

HLA represents each completed chunk as an exact affine state transition and computes content-dependent routing gates from compact self-attentively pooled representatives. Each gate interpolates the chunk's historical transition with the identity map, controlling both additive memory and transformation of earlier states, with effective-support regularization encouraging concentrated routing.

Across Qwen3.5 models from 0.8B to 9B, HLA beats native GDN and fixed chunk mixing by up to 5.57 points on LongBench-V2 and 3.97 on RULER; a from-scratch 1.3B model shows gains growing from 0.83 points at 4K to 4.22 at 32K context.

Original post →

More from Research

Research channel →