MSRA team unveils Universal YOCO: shared-parameters recursion for efficient depth scaling
donglixp · x · 2026-09-11
Researchers at Microsoft Research Asia (Yutao Sun, Li Dong, Furu Wei, et al.) released Universal YOCO (YOCO-U), arXiv:2604.01220. The paper argues standard Transformers scale inference-time compute poorly: looping strategies carry high overhead and KV caches grow with depth. YOCO-U combines the YOCO decoder-decoder architecture with recursion — a Universal Self-Decoder iterates via parameter sharing, confined to shallow efficient-attention layers. YOCO contributes a constant global KV cache and linear prefill, while partial recursion adds representational depth at limited cost. Empirically YOCO-U stays competitive on general and long-context benchmarks. The authors frame it as architecture innovation's 'second half': moving from component innovation to layout innovation, with Looped YOCO completing the stack.
Related event: Microsoft's Universal YOCO enables efficient depth scaling(3 posts)→
More from Research
- World Model RL Debiasing Cuts Cost of Scaling Autonomous Research Agents — illinois · 2026-09-11
- RL-trained agents should carry a strong simulation prior, argues vooooogel — voooooogel · 2026-09-11
- Histed: even if the Navier-Stokes step borrowed human results, AI math progress is real — HistedLab · 2026-09-11
- Humansand launches to simulate humans for RLHF data — gharik · 2026-09-11
- Abstract CoT: latent reasoning cuts tokens up to 11.6x with CoT-level performance — evijit · 2026-09-11
- One person with an AI agent cut Google's quantum ECDSA circuit cost 52%; crowd beat it in 73 hours — anselm · 2026-09-11