DeepSeek V4.1 shrinks cache 437x, Flash beats V4-Pro 39 vs 36 at half the cost
DeepLearningAI · x · 2026-10-08
DeepLearning.AI's The Batch breaks down DeepSeek-V4.1's architecture: since agents read far more context than they write, cache storage became the cost bottleneck, so DeepSeek cut cache to 890 bytes/token — 437x smaller than V1 — while extending input from 4K to 1M tokens with only +25% compute per output token. The Flash variant scores 39 vs 36 for V4-Pro on Artificial Analysis' index at $0.27 vs $0.67/task. Architecture: encoder-decoder MoE with 552B backbone + 196B memory module (8B active reading, 16B generating), first V4 model with image input, 384K max output at 225.6 tokens/s, and halved API prices off-peak.
More from Infra
- WSJ: Broadcom arranging $50B+ financing for OpenAI's custom chips under secret Nexus program — rohanpaul_ai · 2026-10-08
- Reddit asks: have AI scaling laws hit their limit, or is the compute buildout just starting? — StupidDialUp · 2026-10-08
- TRIAGE stabilizes native NVFP4 RL training, hits full-precision quality at 2.3x throughput — InfiX-ai · 2026-10-08
- WSL Containers Now Generally Available: Run Linux Containers Natively on Windows — pavandavuluri · 2026-10-08
- ai& Says It's Japan's Largest Dedicated Inference Provider, Teases Post-Training Offerings — DavidBennett__ · 2026-10-08
- Microsoft unveils Surface Laptop Ultra with Nvidia RTX Spark SoC from $2,599 — Ars Technica AI · 2026-10-08