FULL STORY
DeepSeek V4 Flash: Launch and Surge
DeepSeek launched the V4 Flash public beta, gaining massive traction for its enhanced agentic capabilities and low costs, quickly topping the Ollama growth chart.
2026-08-04 ~ 2026-08-06 · 2 episodes · 11 posts
Episode 1 · DeepSeek V4 Flash Public Beta Launches, Low-Cost Agent Capabilities Draw Attention (2026-08-04, 9 posts)
DeepSeek launched V4 Flash API public beta on Aug 4, featuring significantly enhanced agent capabilities, tool calling, and long-task performance, with extremely low inference cost drawing industry attention. The model has been rolled out on platforms including Together AI and CoreWeave, further driving down market inference pricing with its excellent cost-performance.
Confirmed
- Model specs: 284B total parameters, 13B active per token, 1M token context window, with hybrid attention mechanism.
- In Agent Arena based on 12.5K real agent sessions, DeepSeek-V4-Flash ranked third among open-source models.
- Natively supports Responses API format and Codex adaptation, with significantly enhanced coding abilities in terminal, code repository, and full-stack tasks.
- Together AI has launched the model with adjustable inference modes (low/high/max).
- W&B confirmed DeepSeek-V4-Flash-0731 is now available on CoreWeave's serverless inference platform. Developer @ScottCondron reported excellent performance in data extraction tasks, balancing speed and low cost.
Unconfirmed
- The claim of "1/57 the price of competitors" comes from Together AI's announcement, with unspecified comparison baseline.
- Specific comparison data relayed by @orange, such as "over a billion tokens beating tens of billions of GPT", cannot be verified due to lack of full context in original posts.
Why it matters
- DeepSeek V4 Flash offers frontier-level agent capabilities while drastically reducing inference cost (reportedly 1/57 of competitors).
- According to @teortaxesTex, the Disk KV cache is a killer feature; this architectural breakthrough may seem threatening to other LLM vendors, but due to DeepSeek's fully open-source strategy, it is a huge boon to the entire open-source ecosystem.
- @infwinston and @intellectronica both believe this not only provides developers with a cost-effective choice but also forms a strong market force that effectively drives down AI inference pricing across the industry.
- Together AI puts DeepSeek V4 Flash live, touting stronger coding and 1/57th the price — togethercompute · 2026-08-04
- Together AI launches DeepSeek V4 Flash with adjustable reasoning and cheaper agentic coding — togethercompute · 2026-08-04
- DeepSeek V4 Flash 0731 ships on Together AI with 284B parameters and 1M context — togethercompute · 2026-08-04
- DeepSeek-V4-Flash Lands on Agent Arena at #3 Among Open-Source Models — infwinston · 2026-08-04
- DeepSeek V4 Flash Launches, Driving Down Inference Pricing — intellectronica · 2026-08-05
- DeepSeek-V4-Flash-0731 Goes Live on CoreWeave Serverless Inference — wandb · 2026-08-06
- DeepSeek V4-Flash-0731 Goes Live on CoreWeave's Serverless Inference — _ScottCondron · 2026-08-06
- DeepSeek v4 Flash's Disk KV cache is a killer feature set to reduce industry-wide serving costs — teortaxesTex · 2026-08-06
- Dev Claims DeepSeek v4 flash Delivers Massive Efficiency Gains Over GPT at Half the Cost — oran_ge · 2026-08-06
Episode 2 · DeepSeek-V4-Flash Tops Ollama Growth Chart (2026-08-05, 2 posts)
DeepSeek-V4-Flash has become the fastest-growing model on Ollama. Featuring a 284B MoE architecture, it supports a 1M token context, high inference speeds, and zero data retention.
- DeepSeek-V4-Flash Becomes Fastest Growing Model on Ollama with Zero Data Retention — ollama · 2026-08-05
- DeepSeek-V4-Flash Specs and Benchmarks Revealed: 284B MoE, 1M Token Context — ollama · 2026-08-05