Universal YOCO paper combines recursive compute with efficient attention for depth scaling

donglixp · x · 2026-09-11

A tweet chain points to the arXiv paper 'Universal YOCO for Efficient Depth Scaling' (Yutao Sun, Li Dong, Furu Wei et al.). YOCO-U merges the YOCO decoder-decoder architecture with recursive computation: a parameter-shared Universal Self-Decoder iterates multiple times, confined to shallow efficient-attention layers, giving a constant global KV cache, linear pre-filling, and deeper representations at limited overhead. It stays competitive on general and long-context benchmarks, suggesting efficient-attention + recursion as a promising scaling direction.

Related event: Microsoft's Universal YOCO enables efficient depth scaling(3 posts)→

Original post →

More from Research

Research channel →