Hybrid LMs like Qwen3.5 barely use their recurrent memory; a simple auxiliary pass fixes it

mohitban47 · x · 2026-10-07

New research finds hybrid attention-recurrent LMs (Qwen3.5, Nemotron-H) lean almost entirely on attention: blocking attention collapses accuracy (30%→5%) while blocking recurrence barely hurts (30%→20%), and standard SFT widens the gap.

The fix

Related event: Study: Hybrid LLMs Over-Rely on Attention; Auxiliary Loss Unlocks Recurrence(4 posts)→

Original post →

More from Research

Research channel →