DeepSeek cuts V4-Flash API prices Sept 10, inputs back to pre-hike levels as V4.1 test model appears
机器之心 · wechat · 2026-09-09
DeepSeek announced V4-Flash API price cuts effective Sept 10 at 12:00 Beijing time: off-peak rates of ¥0.02/M tokens (cache hit), ¥1 (cache miss), and ¥4 (output), with peak hours (weekdays 9:00–12:00, 14:00–18:00) at 2x. V4-Pro pricing unchanged.
- Compared to pre-hike prices (Aug 17: 0.02/1/2), inputs return to original levels while output remains 2x; at peak hours it's 2x/2x/4x vs the old flat rate.
- Real-world impact depends on cache hit rate, input/output ratio, and off-peak flexibility — long-context, fixed-system-prompt, and agent-loop workloads benefit most.
- Separately, DeepSeek released a test model deepseek-v4.1-flash-expires-on-0910 at the same billing, native multimodal, faster inference, and lower costs via architecture upgrades — it auto-expires Sept 10, the same day pricing takes effect, hinting at an imminent V4.1 release.
More from Models
- Anthropic Accuses Moonshot of Routing 300K User Queries to Claude via 5,380 Fake Accounts — toptickcrypto · 2026-09-11
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11