DeepSeek cuts flash pricing up to 60% as v4.1 flash quietly goes live

量子位 · wechat · 2026-09-09

Per QbitAI, DeepSeek announced new price cuts for its flash series: per-million-token off-peak cache pricing drops from ¥0.05 to ¥0.02 (-60%), off-peak input from ¥1.5 to ¥1 (-33%), and off-peak output from ¥4.5 to ¥4 (-11%). Peak-hour rates are 2x off-peak.

A new DeepSeek v4.1 flash model has also quietly launched — callable via the model name deepseek-v4.1-flash-expires-on-0910 in the API.

Related event: DeepSeek cuts V4-Flash API prices with new peak/off-peak billing from Sept 10(6 posts)→

Original post →

More from Models

Models channel →