Hybrid LLMs Over-Rely on Attention; Boosting Recurrent Memory Use Gains up to 12.1% on Agentic Tasks

EliasEskin · x · 2026-10-07

A study of recurrent–attention hybrid LMs (e.g., Qwen3.5) finds:

Takeaway: hybrid architectures leave headroom on the table; explicitly steering recurrent memory is a cheap win.

Related event: Hybrid LLMs Over-Rely on Attention; Masking It Boosts Long-Context Performance(2 posts)→

Original post →

More from Research

Research channel →