DeepSeek V4.1-Flash cuts KV cache memory to a quarter, targets cheaper agents
The Decoder · rss · 2026-09-10
- DeepSeek released V4.1-Flash, a multimodal model with 552B total parameters and only 16B active per token.
- The key optimization cuts KV cache memory to 1/4 of its predecessor, aimed at making AI agents far cheaper to run.
- It narrowly beats Opus 5 and GPT-5.6 Sol on the DeepSWE coding benchmark.
- The model ships under the MIT license.
More from Infra
- Autonomous adds Omarchy OS to its $26,100 dual-RTX-5090 AI workstation — dee_hw · 2026-09-10
- 100M output tokens for $60: DeepSeek off-peak pricing undercuts Opus 5 by 40x — airesearch12 · 2026-09-10
- Dev take: token demand will grow far faster than demand for top-line intelligence — willcb · 2026-09-10
- Arm lands Lenovo and ByteDance's Volcengine as first China customers for its AI server chips — pstAsiatech · 2026-09-10
- d-Matrix adopts NVIDIA NVLink Fusion to bring Raptor XPUs to rack-scale deployment — nordicinst · 2026-09-10
- Why DeepSeek might profit despite open weights: it's the only one happy to optimize for its own architecture — yacineMTB · 2026-09-10