Auxiliary loss forces hybrid LMs like Qwen3.5 to actually use recurrent memory, +12.1% on agentic tasks

mohitban47 · x · 2026-10-07

New research finds hybrid LMs (Qwen3.5, Nemotron-H) rely heavily on attention even after SFT, leaving recurrence underused. A simple auxiliary loss forces past information through the recurrent pathway during training, improving QA by 4.6% and agentic tasks by 12.1% on average, with larger gains on longer contexts and up to 28.6% on attention-only models with multiple memory types. Having multiple memory pathways is only half the solution—models must learn to use them.

Related event: Study: Hybrid LLMs Over-Rely on Attention; Auxiliary Loss Unlocks Recurrence(4 posts)→

Original post →

More from Models

Models channel →