Interleaved Head Attention Accepted at NeurIPS, Boosts RULER Retrieval 10-20%
_arohan_ · x · 2026-10-01
Interleaved Head Attention (IHA) has been accepted at NeurIPS 2026. Standard multi-head attention keeps heads isolated, forcing cross-head relation composition across layers; IHA instead builds P pseudo-heads per head (typically P=H), where each pseudo query/key/value is a learned linear combination of all H original ones, enabling within-layer head communication at O(H²P) overhead.
Theory shows parameter savings on synthetic Polynomial (Θ(√k·n²) vs Θ(kn²)) and order-sensitive CPM-3 tasks. On real benchmarks, IHA improves RULER multi-key retrieval by 10-20% (4k-16k) and, after OpenThoughts reasoning fine-tuning, lifts GSM8K by 5.8% and MATH-500 by 2.8% (majority vote).
More from Models
- BuildingBench: GPT-6.1 Sol scores 85.1 at 1/27 the cost of Opus 5.5's 86.8 — ZhitingHu · 2026-10-01
- Gemini 3.8 Flash (high) hits 84.8% on WeirdML v2, first Flash to beat Gemini 3.1 Pro — teortaxesTex · 2026-10-01
- GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support — ycombinator · 2026-10-01
- Researchers extract raw reasoning traces again, now from GPT-6 Astra — maksym_andr · 2026-10-01
- Fireworks launches Ember-1, built on Kimi K3, matching quality with fewer reasoning tokens — sophiamyang · 2026-10-01
- Anthropic's AI cracks decades-old percolation conjecture, called Fields Medal-level work — Dr_Singularity · 2026-10-01