CUHK study maps when recurrence helps in looped language models, proposes history-state injection
CUHK-CSE · hf · 2026-10-01
A CUHK study systematically examines when recurrence is effective in looped language models (LoopLMs), which add computational depth via parameter sharing without extra parameters.
Key findings
- Recurrence improves reasoning beyond the training horizon but degrades knowledge performance; harder instances don't consistently benefit more.
- Performance depends on how layers and recurrent iterations are allocated — effective depth alone can't predict behavior; non-recurrent output layers improve robustness to under-unrolling.
- Conventional initial-state injection offers limited robustness; the proposed channel-wise history-state injection with timestep conditioning better preserves knowledge and stays robust across inference budgets.
The paper offers practical design guidelines for LoopLMs under variable inference budgets.
More from Models
- Rumor: Gemini 4 spotted with SOTA knowledge scores, coding near Astra/Fable 5.1 level — haider1 · 2026-10-01
- GPT-6.1 Sol benchmarks across all effort levels land, making GPT-6 Astra hard to justify — PawelHuryn · 2026-10-01
- Do uncensored open-weight models actually matter? Reddit sparks debate — EmilPi · 2026-10-01
- Reddit user claims Gemini 4 Argon "solved hallucinations" — and nobody's talking about it — drhenriquesoares · 2026-10-01
- Gemini element-naming race heats up: Neon and Argon taken, Krypton is next as Google rejoins the frontier — tkipf · 2026-10-01
- Rumor: Gemini 4 Argon will launch as Ultra subscriber exclusive at first — opmgyhx · 2026-10-01