Universal YOCO paper combines recursive compute with efficient attention for depth scaling
donglixp · x · 2026-09-11
A tweet chain points to the arXiv paper 'Universal YOCO for Efficient Depth Scaling' (Yutao Sun, Li Dong, Furu Wei et al.). YOCO-U merges the YOCO decoder-decoder architecture with recursive computation: a parameter-shared Universal Self-Decoder iterates multiple times, confined to shallow efficient-attention layers, giving a constant global KV cache, linear pre-filling, and deeper representations at limited overhead. It stays competitive on general and long-context benchmarks, suggesting efficient-attention + recursion as a promising scaling direction.
Related event: Microsoft's Universal YOCO enables efficient depth scaling(3 posts)→
More from Research
- Mathathon organizers respond to open letter, pledge a 'responsible AI math' experiment — soumitrashukla9 · 2026-09-11
- DARPA launches expMath program; researchers share AI-driven mathematical discovery work — wellecks · 2026-09-11
- VidMap: ETH researchers open-source video Structure-from-Motion system, ECCV 2026 — rsasaki0109 · 2026-09-11
- AI-agent eye clinic in China among first real-world AI-native healthcare deployments — EricTopol · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Caltech undergrads defend AI math contest Mathathon against open letter criticism — Singularitarian · 2026-09-11