antirez: DeepSeek's shared KV cache architecture enables fast big prefills on low-memory local setups
antirez · x · 2026-09-11
antirez concludes that DeepSeek v4.1 Flash's encoder/decoder architecture with shared KV cache not only removes compute during prefill but also enables very fast large prefills in low-memory local inference setups. He notes a full resident decoder trashes the experts cache, but the new SSD streaming implementation hides loading well and compensates within seconds, provided the machine has decent switch speeds.
More from Infra
- 176 KB C Program Runs 2.78T-Param Kimi K3 on a Single CPU with 8.24 GB RAM — techNmak · 2026-09-11
- Blackstone's Biggest AI Bet Is Compute, Backing Deals with Google, Nvidia, Anthropic — abhiadesai · 2026-09-11
- Frontier models now independently reach for speculative decoding and kernel optimization on InferenceBench — maksym_andr · 2026-09-11
- A Beginner-Friendly Guide to Budget Multi-GPU Local LLM Setups — lblblllb · 2026-09-11
- Chinese Nvidia challenger Enflame jumps 179% in Shanghai debut, raises $910M — pstAsiatech · 2026-09-11
- Qwen3.8 Flash Next hits 49 tok/s locally on 2x RTX 3090 with FlashNext llama.cpp fork — whiteh4cker · 2026-09-11