YOCO explained: one shared KV cache reused across the model's second half
stochasticchasm · x · 2026-09-11
On the YOCO (You Only Cache Once) architecture, the author clarifies the design is quite simple:
- The first half of the model is standard;
- Hidden states are projected once more to K/Vs, and that identical KV cache is reused by every layer in the second half, slashing cache overhead.
The author also notes the work does acknowledge YoCo, and quips the naming should be YOCO, not YoCo.
Related event: Inside YOCO: later layers share a single KV cache(4 posts)→
More from Research
- Another Model Shifts to Muon Optimizer as It Emerges as the Training Default — stochasticchasm · 2026-09-11
- SkillAdam ports Adam's moment estimates to agent skill docs to fix self-evolution loops — dair_ai · 2026-09-11
- Researcher: 95% confidence intervals may really cover just 10-25% of the truth — RexDouglass · 2026-09-11
- Crucible: A Transformer-Free Neurosymbolic Stack Pairing Mamba2 with Z3 Verification — JustRoccat · 2026-09-11
- 95% Confidence Intervals Really Cover Only 10-25% of the Time, Statistician Notes — RexDouglass · 2026-09-11
- Falsifiability debate: does AI doom make testable claims about present evidence — lsindjowt · 2026-09-11