Masked-attention fine-tuning fixes hybrid LLMs ignoring recurrence, +4.6% QA, +12.1% agentic

mohitban47 · x · 2026-10-07

Next-gen hybrid LLMs (e.g. Qwen-3.5) mix recurrent layers with attention, but researchers found naive SFT makes them rely almost entirely on attention and ignore the recurrent pathway — even though attention excels at needle-in-a-haystack retrieval while recurrence excels at aggregating information spread across long contexts.

Their fix: during fine-tuning, run a second forward pass with attention layers masked so outputs can't attend to earlier context while recurrent layers still see the full sequence, applying standard next-token loss to force use of the recurrent state.

Results: +4.6% average on long-context QA, +12.1% on agentic tasks, and +28.6% on attention-only models with multiple memory types, generalizing across multiple hybrid LLMs.

Related event: Study: Hybrid LLMs Over-Rely on Attention; Auxiliary Loss Unlocks Recurrence(4 posts)→

Original post →

More from Research

Research channel →