MSRA team unveils Universal YOCO: shared-parameters recursion for efficient depth scaling

donglixp · x · 2026-09-11

Researchers at Microsoft Research Asia (Yutao Sun, Li Dong, Furu Wei, et al.) released Universal YOCO (YOCO-U), arXiv:2604.01220. The paper argues standard Transformers scale inference-time compute poorly: looping strategies carry high overhead and KV caches grow with depth. YOCO-U combines the YOCO decoder-decoder architecture with recursion — a Universal Self-Decoder iterates via parameter sharing, confined to shallow efficient-attention layers. YOCO contributes a constant global KV cache and linear prefill, while partial recursion adds representational depth at limited cost. Empirically YOCO-U stays competitive on general and long-context benchmarks. The authors frame it as architecture innovation's 'second half': moving from component innovation to layout innovation, with Looped YOCO completing the stack.

Related event: Microsoft's Universal YOCO enables efficient depth scaling(3 posts)→

Original post →

More from Research

Research channel →