DeepSeek-V4-Flash Official API Launches with Enhanced Agent Capabilities and Ultra-low Pricing

智东西 · wechat · 2026-07-31

DeepSeek-V4-Flash official API has entered public beta, with the Pro version slated for early August. The Flash model utilizes a MoE architecture (284B total / 13B active params), supports 1M token context, and natively supports the OpenAI Responses API format.

The update focuses on significantly enhanced agentic capabilities, outperforming the April preview version across multiple benchmarks like TerminalBench and Cybergym. It also introduces aggressive long-context optimizations, drastically reducing inference compute and KV Cache footprint for 1M token scenarios.

Pricing remains highly competitive: Flash cached input costs 0.2 RMB/M tokens and output 2 RMB/M tokens; Pro output is 24 RMB/M tokens. The article notes that as top tier competitors like OpenAI also cut prices, low API costs are no longer the sole differentiator; the true battleground has shifted to practical agent performance.

Related event: DeepSeek-V4-Flash Official API Launches with Major Agent Upgrades(18 posts)→

Original post →

More from coding & agent

coding & agent channel →