Hybrid LLMs Over-Rely on Attention; Boosting Recurrent Memory Use Gains up to 12.1% on Agentic Tasks
EliasEskin · x · 2026-10-07
A study of recurrent–attention hybrid LMs (e.g., Qwen3.5) finds:
- Problem: despite having two complementary memory pathways—attention and recurrence—hybrid models rely heavily on attention, even after SFT, leaving recurrence underused
- Fix: training interventions that encourage recurrent memory use
- Results: biggest gains on long-context tasks—QA up 4.6% on average, agentic tasks up 12.1%, and up to 28.6% on attention-only models with multiple memory types
- Findings generalize across multiple hybrid LLMs
Takeaway: hybrid architectures leave headroom on the table; explicitly steering recurrent memory is a cheap win.
More from Research
- Paradigm releases Limite 1B -Violetto technical report: a tiny model trained from scratch to max out math per parameter — tensorqt · 2026-10-07
- AFP-GIC Cuts Generative Image Codec Latency 18.1% and Params 20.5% vs DC-VIC — SantaClaraUniversity · 2026-10-07
- Microsoft: LLMs Are Already Jev-Style Decision Models, Fine-Tuning Isn't Always Needed — microsoft · 2026-10-07
- APO Enables Personalized LLM Alignment with Only 20 Local User Examples — Liyan Yang · 2026-10-07
- Priced Guidance: quantifying whether LLMs can generate novel research ideas via compression — stanfordnlp · 2026-10-07
- Stratego, the Hidden-Information Game That Stumped AI, Finally Falls — ArtificialOther · 2026-10-07