Hybrid LMs like Qwen3.5 barely use their recurrent memory; a simple auxiliary pass fixes it
mohitban47 · x · 2026-10-07
New research finds hybrid attention-recurrent LMs (Qwen3.5, Nemotron-H) lean almost entirely on attention: blocking attention collapses accuracy (30%→5%) while blocking recurrence barely hurts (30%→20%), and standard SFT widens the gap.
The fix
- An auxiliary pass where attention can't see the past, forcing reliance on recurrent memory
- Overall gains: +4.6% on QA, +12.1% on agentic tasks, up to +28.6% for attention-only models with extra memory types
- Largest benefits on long contexts, multi-evidence questions, and long-horizon agent tasks
More from Research
- GroundedSLAM debuts, decisively beating all methods on Meta's egocentric SLAM benchmark — Scobleizer · 2026-10-07
- HCI researcher begs authors to stop claiming 'reflexive' thematic analysis without reflexivity — IanArawjo · 2026-10-07
- Reza Zadeh claims faster matrix multiplication algorithm, suspects labs near exponent 2 — Reza_Zadeh · 2026-10-07
- Redditor proposes graph-based deterministic modeling to make LLM finance agents trustworthy — jonnylegs · 2026-10-07
- COLM 2026 poster presents scaling test-time compute for agentic coding — dan_fried · 2026-10-07
- COLM 2026 Efficient Reasoning workshop lands Friday, with panel featuring top researchers — tydsh · 2026-10-07