Auxiliary loss forces hybrid LMs like Qwen3.5 to actually use recurrent memory, +12.1% on agentic tasks
mohitban47 · x · 2026-10-07
New research finds hybrid LMs (Qwen3.5, Nemotron-H) rely heavily on attention even after SFT, leaving recurrence underused. A simple auxiliary loss forces past information through the recurrent pathway during training, improving QA by 4.6% and agentic tasks by 12.1% on average, with larger gains on longer contexts and up to 28.6% on attention-only models with multiple memory types. Having multiple memory pathways is only half the solution—models must learn to use them.
More from Models
- Early verdict on Meta Muse: "not very good" — HanchungLee · 2026-10-07
- Claude Surprises User by Offering to Switch to Work Mode Mid-Maxscript Coding — Efistoffeles · 2026-10-07
- DIY eval: pplx-decider-1.1 hits 95.6% agreement with a frontier model — bo_wangbo · 2026-10-07
- Google's EmbeddingGemma 2 (740M) claims to beat embedding rivals twice its size — The Decoder · 2026-10-07
- Mistral Large 4 launches as CEO claims it beats Chinese models on cyber capabilities — BLUECOW009 · 2026-10-07
- Florida woman charged with felony after Claude flagged her threat and a human reviewer tipped police — XFreeze · 2026-10-07