FULL STORY
DeepSeek V4.1 Flash Limited Beta and New Pricing Revealed
DeepSeek opened a limited beta for V4.1 Flash featuring a new architecture and native multimodality, with reports of steep price cuts across the Flash series starting September 10.
2026-09-08 ~ 2026-09-09 · 3 episodes · 16 posts
Episode 1 · DeepSeek V4.1 Flash Opens Limited Beta with New Architecture and Native Multimodality (2026-09-08, 12 posts)
On September 8, multiple sources revealed that DeepSeek V4.1 Flash has entered internal testing as an "interim version," with core upgrades including a brand-new model architecture and native multimodal support. Officially, it is billed as more capable, faster, and cheaper.
Confirmed
- Synced reports that DeepSeek officially announced the V4.1 Flash interim version for internal testing, with major upgrades including a new model architecture, native multimodality, and stronger capabilities at higher speed and lower cost.
- How to use it: keep the baseurl unchanged and set the model name to an ID starting with deepseek-v4.1-flash-expires-on-09.
- According to an official WeChat group notice relayed by vista8, billing is the same as V4 Flash, with up to 20 concurrent requests per account.
Unconfirmed
- The news initially broke via bloggers such as op7418 and AGI Hunt from DeepSeek's official WeChat group, all stressing it lacked confirmation from an official public announcement; Synced later reported it as "officially announced," but a formal public announcement and full capability details are still pending.
- teortaxesTex relayed a community joke hoping it can at least beat GLM 5.3 Flash beyond vision—this is only community opinion with no benchmark evidence.
Why it matters
- This is the first time DeepSeek has been reported to release an internal test as an "interim version (expires-on)," with billing and concurrency rules clarified in tandem, reflecting a product cadence of iterating through real developer testing before official release.
- Native multimodality marks a major upgrade to DeepSeek's flagship model roadmap, and the community has already started benchmarking it against competitors like GLM 5.3 Flash—worth watching for real-world performance once the official version ships.
- How to call DeepSeek V4.1 Flash beta: model name and rate limits revealed — 赛博禅心 · 2026-09-08
- Rumored DeepSeek V4.1 Flash surfaces, user hopes it beats GLM 5.3 Flash — teortaxesTex · 2026-09-08
- DeepSeek V4.1 Flash mid-cycle build enters beta with native multimodal support — AGI Hunt · 2026-09-08
- Rumor: DeepSeek V4.1 Flash released with stronger native multimodal — op7418 · 2026-09-08
- DeepSeek V4.1 Flash reportedly in private beta with new architecture and native multimodal support — vista8 · 2026-09-08
- DeepSeek opens internal beta of V4.1 Flash with native multimodal support — jiqizhixin · 2026-09-08
- DeepSeek spotted testing V4-Flash-Vision: new arch, native multimodal, same price — teortaxesTex · 2026-09-08
- DeepSeek quietly tests V4.1 Flash API beta with new architecture and native multimodality — tokenbender · 2026-09-08
- DeepSeek Opens Limited-Time Beta of V4.1Flash: New Architecture, Native Multimodal, Expires Sept 10 — APPSO · 2026-09-08
- DeepSeek V4.1 Flash interim build reportedly in beta: native multimodal, faster and cheaper — aigclink · 2026-09-08
- DeepSeek Flash 4.1 spotted testing via API with new architecture, release imminent — kimmonismus · 2026-09-08
- DeepSeek V4.1 Flash in internal beta: native multimodal, same price as V4 Flash — Nunki08 · 2026-09-08
Episode 2 · DeepSeek Opens V4.1 Flash Beta with Aggressive Pricing (2026-09-08, 2 posts)
DeepSeek has opened beta access to V4.1 Flash, featuring a new architecture, native multimodality, and outputs up to 507 tokens/s. It also launched a two-day promotion where $1 buys roughly 58 million tokens.
- DeepSeek V4.1 Flash beta goes live with new architecture, native multimodality, up to 507 tokens/s — 智东西 · 2026-09-08
- DeepSeek v4.1 Flash flash sale: 58M tokens for $1, available for 2 days only — MicahBerkley · 2026-09-09
Episode 3 · DeepSeek's new pricing leaks flash discounts, off-peak cache at 0.02 yuan (2026-09-08, 2 posts)
Leaked details indicate DeepSeek will roll out new pricing on September 10, cutting flash-series input to 1 yuan per million tokens off-peak with cache hits as low as 0.02 yuan, while raising output prices.
- DeepSeek cuts flash-series prices to 1 yuan/M input tokens, launches v4.1-flash — 赛博禅心 · 2026-09-08
- DeepSeek's Post-Sept 10 Pricing: Premium Output, Off-Peak Cache Hits at ¥0.02 — teortaxesTex · 2026-09-08