FULL STORY

DeepSeek V4 Flash: Launch and Surge

DeepSeek launched the V4 Flash public beta, gaining massive traction for its enhanced agentic capabilities and low costs, quickly topping the Ollama growth chart.

2026-08-04 ~ 2026-08-06 · 2 episodes · 11 posts

Episode 1 · DeepSeek V4 Flash Public Beta Launches, Low-Cost Agent Capabilities Draw Attention (2026-08-04, 9 posts)

DeepSeek launched V4 Flash API public beta on Aug 4, featuring significantly enhanced agent capabilities, tool calling, and long-task performance, with extremely low inference cost drawing industry attention. The model has been rolled out on platforms including Together AI and CoreWeave, further driving down market inference pricing with its excellent cost-performance.

Confirmed

  • Model specs: 284B total parameters, 13B active per token, 1M token context window, with hybrid attention mechanism.
  • In Agent Arena based on 12.5K real agent sessions, DeepSeek-V4-Flash ranked third among open-source models.
  • Natively supports Responses API format and Codex adaptation, with significantly enhanced coding abilities in terminal, code repository, and full-stack tasks.
  • Together AI has launched the model with adjustable inference modes (low/high/max).
  • W&B confirmed DeepSeek-V4-Flash-0731 is now available on CoreWeave's serverless inference platform. Developer @ScottCondron reported excellent performance in data extraction tasks, balancing speed and low cost.

Unconfirmed

  • The claim of "1/57 the price of competitors" comes from Together AI's announcement, with unspecified comparison baseline.
  • Specific comparison data relayed by @orange, such as "over a billion tokens beating tens of billions of GPT", cannot be verified due to lack of full context in original posts.

Why it matters

  • DeepSeek V4 Flash offers frontier-level agent capabilities while drastically reducing inference cost (reportedly 1/57 of competitors).
  • According to @teortaxesTex, the Disk KV cache is a killer feature; this architectural breakthrough may seem threatening to other LLM vendors, but due to DeepSeek's fully open-source strategy, it is a huge boon to the entire open-source ecosystem.
  • @infwinston and @intellectronica both believe this not only provides developers with a cost-effective choice but also forms a strong market force that effectively drives down AI inference pricing across the industry.

Episode 2 · DeepSeek-V4-Flash Tops Ollama Growth Chart (2026-08-05, 2 posts)

DeepSeek-V4-Flash has become the fastest-growing model on Ollama. Featuring a 284B MoE architecture, it supports a 1M token context, high inference speeds, and zero data retention.