DeepSeek V4.1 Flash cuts global KV cache to 890 bytes/token, but HBM demand may rise with agent swarms

teortaxesTex · x · 2026-09-12

An analysis of DeepSeek V4.1 Flash argues HBM capacity is less of a bottleneck: via CSA2 (cross-layer global KV reuse with FP4 KV caching and Top-K indices) plus SWA Bounded Replay, global KV drops from 3.5KB/tok to 890 bytes and persistent KV cache footprint falls to 1/8 of V4 Flash. But a rebuttal notes HBM demand isn't going away — at 4 agents, V4.1's HBM consumption/second exceeds V4's, and at 64 it exceeds V3.2. Jevons paradox strikes again.

Related event: DeepSeek V4.1 Flash Shrinks KV Cache 437x, Sparking HBM Debate(2 posts)→

Original post →

More from Infra

Infra channel →