FULL STORY

DeepSeek V4.1 Flash Limited Beta and New Pricing Revealed

DeepSeek opened a limited beta for V4.1 Flash featuring a new architecture and native multimodality, with reports of steep price cuts across the Flash series starting September 10.

2026-09-08 ~ 2026-09-09 · 3 episodes · 16 posts

Episode 1 · DeepSeek V4.1 Flash Opens Limited Beta with New Architecture and Native Multimodality (2026-09-08, 12 posts)

On September 8, multiple sources revealed that DeepSeek V4.1 Flash has entered internal testing as an "interim version," with core upgrades including a brand-new model architecture and native multimodal support. Officially, it is billed as more capable, faster, and cheaper.

Confirmed

  • Synced reports that DeepSeek officially announced the V4.1 Flash interim version for internal testing, with major upgrades including a new model architecture, native multimodality, and stronger capabilities at higher speed and lower cost.
  • How to use it: keep the baseurl unchanged and set the model name to an ID starting with deepseek-v4.1-flash-expires-on-09.
  • According to an official WeChat group notice relayed by vista8, billing is the same as V4 Flash, with up to 20 concurrent requests per account.

Unconfirmed

  • The news initially broke via bloggers such as op7418 and AGI Hunt from DeepSeek's official WeChat group, all stressing it lacked confirmation from an official public announcement; Synced later reported it as "officially announced," but a formal public announcement and full capability details are still pending.
  • teortaxesTex relayed a community joke hoping it can at least beat GLM 5.3 Flash beyond vision—this is only community opinion with no benchmark evidence.

Why it matters

  • This is the first time DeepSeek has been reported to release an internal test as an "interim version (expires-on)," with billing and concurrency rules clarified in tandem, reflecting a product cadence of iterating through real developer testing before official release.
  • Native multimodality marks a major upgrade to DeepSeek's flagship model roadmap, and the community has already started benchmarking it against competitors like GLM 5.3 Flash—worth watching for real-world performance once the official version ships.

Episode 2 · DeepSeek Opens V4.1 Flash Beta with Aggressive Pricing (2026-09-08, 2 posts)

DeepSeek has opened beta access to V4.1 Flash, featuring a new architecture, native multimodality, and outputs up to 507 tokens/s. It also launched a two-day promotion where $1 buys roughly 58 million tokens.

Episode 3 · DeepSeek's new pricing leaks flash discounts, off-peak cache at 0.02 yuan (2026-09-08, 2 posts)

Leaked details indicate DeepSeek will roll out new pricing on September 10, cutting flash-series input to 1 yuan per million tokens off-peak with cache hits as low as 0.02 yuan, while raising output prices.