DeepSeek V4.1 Flash Is an Architecture Reset: KV Cache Cut to 1/4 by Deleting Its Own Ideas
teortaxesTex · x · 2026-09-16
A Zhihu breakdown by Exhalation argues DeepSeek's V4.1 Flash is an architecture reset, not just a cheaper model: despite more total and active parameters than its predecessor, working KV cache drops to one quarter and persistent cache storage to one eighth. The gains come from deleting old modules, sharing KV states, and recomputing local context.
- MTP removed: external draft models like DSpark eroded its speculative-decoding value, and the auxiliary loss no longer justified its memory cost
- Heavily Compressed Attention removed: ambiguous global-summary role, hard to pair with FP4 storage
- Dense warmup dropped in favor of a sparser training path
teortaxesTex frames this as DeepSeek pushing the quadratic component into ever rarer operations—accepting the bitter lesson even in its own science, searching for the irreducible part.
More from Infra
- Microsoft bets on local AI: Windows agent stack spans $800 Copilot+ PCs to 1T-param DGX Station — ryanshrout · 2026-09-16
- PlanetScale Traffic Control lets you budget DB resources per app name — DanielLockyer · 2026-09-16
- GPUs as VC value-add: European AI startups' top constraint is compute access — nellimorgulchik · 2026-09-16
- Community squeezes a 124B model onto a 128GB DGX Spark with quantization and kernel fixes — alifcoder · 2026-09-16
- One chart explains how CPU, GPU and TPU differ — mdancho84 · 2026-09-16
- Even 4-year-old GPUs are repricing up: 20% renewal premium, contracts up 125% — tengyanAI · 2026-09-16