antirez: DeepSeek's shared KV cache architecture enables fast big prefills on low-memory local setups

antirez · x · 2026-09-11

antirez concludes that DeepSeek v4.1 Flash's encoder/decoder architecture with shared KV cache not only removes compute during prefill but also enables very fast large prefills in low-memory local inference setups. He notes a full resident decoder trashes the experts cache, but the new SSD streaming implementation hides loading well and compensates within seconds, provided the machine has decent switch speeds.

Related event: antirez runs DeepSeek v4.1 Flash locally on Macs with SSD streaming and RDMA tensor parallel(10 posts)→

Original post →

More from Infra

Infra channel →