DeepSeek-V4.1-Flash Hits the API: 1/8 the SSD Cache, V4-Pro Being Phased Out
deepseek_ai · x · 2026-09-10
DeepSeek announced V4.1-Flash is now live on the DeepSeek API (model name deepseek-flash) with native multimodal support, plus migration and cost-cutting details:
- Old models retired: V4-Flash and V4-Flash-Vision-Exp are retired; legacy model names temporarily route to V4.1-Flash
- Independent tests put V4.1-Flash ahead of flagship V4-Pro on performance, cost, speed, and total runtime — V4-Pro is being phased out. From Sept 14, 04:00 UTC 2026, all deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches
- Massive KV cache compression: just 1/4 the HBM and 1/8 the SSD storage vs the previous generation. DeepSeek notes cache-hit charges are a large share of agent costs, so compression cuts those significantly
- Official partners WorkBuddyAI (including Codebuddy) and opencode fully support V4.1-Flash
Related event: DeepSeek Launches V4.1 Flash on API, Set to Retire V4 Pro(5 posts)→
More from Infra
- vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture — vllm_project · 2026-09-10
- No one matches its inference economics; commentator suggests 10-20% cut from infra providers — zephyr_z9 · 2026-09-10
- Apple A20 Pro Neural Engine projected to hit 140+ TFLOPS, a third of an A100 for local inference — AIFlow_ML · 2026-09-10
- Inside V4.1 Flash: Per-Token KV Cache Down to 890 Bytes, Peak-Valley Pricing — 赛博禅心 · 2026-09-10
- V4.1 Flash KV cache is 437x smaller than V1, easing China's memory bottleneck — ChrisGPT · 2026-09-10
- Aspen Digital webinar reimagines data centers as public compute for communities — lfschiavo · 2026-09-10