DeepSeek V4.1 Flash cuts global KV cache to 890 bytes/token, but HBM demand may rise with agent swarms
teortaxesTex · x · 2026-09-12
An analysis of DeepSeek V4.1 Flash argues HBM capacity is less of a bottleneck: via CSA2 (cross-layer global KV reuse with FP4 KV caching and Top-K indices) plus SWA Bounded Replay, global KV drops from 3.5KB/tok to 890 bytes and persistent KV cache footprint falls to 1/8 of V4 Flash. But a rebuttal notes HBM demand isn't going away — at 4 agents, V4.1's HBM consumption/second exceeds V4's, and at 64 it exceeds V3.2. Jevons paradox strikes again.
Related event: DeepSeek V4.1 Flash Shrinks KV Cache 437x, Sparking HBM Debate(2 posts)→
More from Infra
- The Global Race for Cheap Power: Where AI Data Centers Should Actually Go — pravchaw · 2026-09-12
- VCs float 'hardware revenue derivative': fund compute costs via revenue share, not equity — ns123abc · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12
- Yutori's Navigator n2 runs browser agents at $1.46 per task on OSWorld 2.0 vs $13-$40+ for frontier models — DhruvBatra_ · 2026-09-12
- Auto-derived FlashAttention with SMEM and tensor core assignment shown off — vtabbott_ · 2026-09-12
- AI progress timing debate: same-node hardware gains deliver a one-time compute windfall — cis_female · 2026-09-12