Masked-attention fine-tuning fixes hybrid LLMs ignoring recurrence, +4.6% QA, +12.1% agentic
mohitban47 · x · 2026-10-07
Next-gen hybrid LLMs (e.g. Qwen-3.5) mix recurrent layers with attention, but researchers found naive SFT makes them rely almost entirely on attention and ignore the recurrent pathway — even though attention excels at needle-in-a-haystack retrieval while recurrence excels at aggregating information spread across long contexts.
Their fix: during fine-tuning, run a second forward pass with attention layers masked so outputs can't attend to earlier context while recurrent layers still see the full sequence, applying standard next-token loss to force use of the recurrent state.
Results: +4.6% average on long-context QA, +12.1% on agentic tasks, and +28.6% on attention-only models with multiple memory types, generalizing across multiple hybrid LLMs.
More from Research
- Precinct6 cybersecurity dataset (1M-10M rows, MITRE ATT&CK) trends on Hugging Face — witfoo · 2026-10-07
- Guidance-TTT trains an 8B model at test time, beats best results on 4 discovery tasks — pliang279 · 2026-10-07
- Workhorse: whole-body loco-manipulation from human demos, no teleop, one take fully autonomous — ChongZzZhang · 2026-10-07
- S2PD: serial computation in high-noise diffusion makes video models follow physics — elliottszwu · 2026-10-07
- Tencent compresses 770B Hunyuan Hy4 Preview into 214 GiB at ~2.38 bits per weight — OpeningRock8761 · 2026-10-07
- DIY eval: pplx-decider-1.1 hits 95.6% agreement with a frontier model — bo_wangbo · 2026-10-07