DeepSeek cuts flash-series prices to 1 yuan/M input tokens, launches v4.1-flash
赛博禅心 · wechat · 2026-09-08
DeepSeek will significantly cut flash-series pricing effective Sept 10, 2026, 12:00 CST (yuan per million tokens, pending official confirmation): off-peak input 1.5→1, off-peak output 4.5→4, off-peak cache 0.05→0.02. Peak-hour prices are 2x off-peak.
A new model, deepseek-v4.1-flash, is also live: same baseurl, callable via model name deepseek-v4.1-flash-expires-on-0910, billed the same as deepseek-v4-flash with a 20-concurrent-request cap per account.
Related event: DeepSeek's new pricing leaks flash discounts, off-peak cache at 0.02 yuan(2 posts)→
More from Models
- OpenAI burned millions in compute on $1M prize: 10,000 agents, 88 hours, 130B tokens — Hesamation · 2026-09-09
- 4.9M messages, 300B tokens, 10,000 agents: details of OpenAI's Navier–Stokes run — willdepue · 2026-09-09
- Qwen quietly releases Drive-1.0-4B, a 4B driving model finetuned from Qwen3.5 — FullstackSensei · 2026-09-09
- Skeptic mocks OpenAI's containment claim: couldn't even manage 1,000 instances, now 10,000 — scaling01 · 2026-09-09
- Gary Marcus mocks GPT-6 Astra video: OpenAI's 'human-level intelligence' claim is a lie — GaryMarcus · 2026-09-09
- OpenAI details Navier-Stokes proof: a spiraling, elongating vortex forms a finite-time singularity, formalized in Lean — OpenAI · 2026-09-09