DeepSeek V4.1 shrinks cache 437x, Flash beats V4-Pro 39 vs 36 at half the cost

DeepLearningAI · x · 2026-10-08

DeepLearning.AI's The Batch breaks down DeepSeek-V4.1's architecture: since agents read far more context than they write, cache storage became the cost bottleneck, so DeepSeek cut cache to 890 bytes/token — 437x smaller than V1 — while extending input from 4K to 1M tokens with only +25% compute per output token. The Flash variant scores 39 vs 36 for V4-Pro on Artificial Analysis' index at $0.27 vs $0.67/task. Architecture: encoder-decoder MoE with 552B backbone + 196B memory module (8B active reading, 16B generating), first V4 model with image input, 384K max output at 225.6 tokens/s, and halved API prices off-peak.

Original post →

More from Infra

Infra channel →