DeepSeek V4.1 Flash Shrinks KV Cache 437x, Sparking HBM Debate

DeepSeek V4.1 Flash compresses KV cache to 890 bytes per token, 437x smaller than V1, easing HBM capacity constraints for China's AI compute. Analysts now debate whether HBM remains a hard requirement.

2026-09-10 ~ 2026-09-12 · 2 related posts