Inside V4.1 Flash: Per-Token KV Cache Down to 890 Bytes, Peak-Valley Pricing
赛博禅心 · wechat · 2026-09-10
DeepSeek V4.1 Flash is open-sourced: a 550B-parameter asymmetric MoE (8B active on input, 16B on output) focused on agent tasks—DeepSWE v1.1 up from 62.7 to 74.2, Terminal-Bench 3.0 from 11.8 to 30.0, beating Kimi K3, GLM 5.3, Opus 5, and GPT 5.6-Sol on several benchmarks, though GPQA Diamond still trails V4 Pro.
Per-token KV Cache drops from 3,514 bytes to 890 (about 1/437 of V1), with cache-related HBM needs at 1/4 and SSD at 1/8. API pricing is peak-valley: off-peak cache-hit input ¥0.02, miss ¥1, output ¥4 per million tokens; peak is 2x. Old model names will route to the new model.
The companion DeepSeek Harness v0.1.5 is deeply co-trained with the model: file uploads/preview, bidirectional parent-child agent communication, and experimental Agent Teams. Community cases include a paper-cut game built in two prompts. Teams with 2,000 GPUs can contact DeepSeek for large-scale deployment.
More from Infra
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested — No_Run8812 · 2026-09-10
- Screenshot surfaces rare admission of 72-hour KV cache limits in V4-era architecture — zephyr_z9 · 2026-09-10
- DeepSeek cut per-token KV cache size by 54x in nine months — zephyr_z9 · 2026-09-10
- vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture — vllm_project · 2026-09-10