Interleaved Head Attention Accepted at NeurIPS, Boosts RULER Retrieval 10-20%

_arohan_ · x · 2026-10-01

Interleaved Head Attention (IHA) has been accepted at NeurIPS 2026. Standard multi-head attention keeps heads isolated, forcing cross-head relation composition across layers; IHA instead builds P pseudo-heads per head (typically P=H), where each pseudo query/key/value is a learned linear combination of all H original ones, enabling within-layer head communication at O(H²P) overhead.

Theory shows parameter savings on synthetic Polynomial (Θ(√k·n²) vs Θ(kn²)) and order-sensitive CPM-3 tasks. On real benchmarks, IHA improves RULER multi-key retrieval by 10-20% (4k-16k) and, after OpenThoughts reasoning fine-tuning, lifts GSM8K by 5.8% and MATH-500 by 2.8% (majority vote).

Original post →

More from Models

Models channel →