DeepSeek V4.1 Flash reportedly adopts YOCO architecture for order-of-magnitude KV savings

donglixp · x · 2026-09-11

DeepSeek V4.1 Flash has reportedly embraced YOCO (You Only Cache Once) as the soul of its architecture, per a credible unconfirmed leak, inheriting YOCO's order-of-magnitude advantages in prefill and KV caching. The decoder-decoder architecture, from a 2024 paper, caches KV pairs only once while retaining global attention, extends to 1M context with near-perfect needle retrieval, and improves inference memory, prefill latency, and throughput by orders of magnitude.

Related event: Inside DeepSeek V4.1 Flash: YOCO at its core and KV cache reuse(7 posts)→

Original post →

More from Models

Models channel →