FULL STORY

DeepSeek V4 Flash: Open-Source Disruption and Tests

DeepSeek V4 Flash sparked a wave of testing and quickly topped platform usage charts, offering top-tier performance at a fraction of competitors' costs.

2026-07-31 ~ 2026-08-08 · 6 episodes · 168 posts

Episode 1 · DeepSeek V4 Flash Hands-on: Low Cost, High Performance Sparks Community Buzz (2026-07-31, 33 posts)

DeepSeek-V4-Flash-0731 has sparked extensive community testing due to its ultra-low cost and near-top-tier performance. Benchmarks show a 50 intelligence score, matching Gemini 3.6 Flash, with real-world capabilities close to Claude 3.5 Sonnet. In multiple complex task tests, it demonstrated exceptional cost-effectiveness in the AI price war.

Confirmed

  • Performance benchmarks: According to Artificial Analysis, DeepSeek V4 Flash (0731) scored 50, matching Gemini 3.6 Flash and approaching the record 51 set in March 2026. @bindureddy noted its real-world capability is close to Claude 3.5 Sonnet.
  • Cost advantage: @teortaxesTex pointed out that while GPT-5.6 Luna has low input token prices, its output token price ($1.20/M) is much higher than DeepSeek's ($0.28/M). @autooff estimated that for 7 billion tokens per month, Luna would cost $250, while DeepSeek would cost only a third. @zainhas's comparison showed per-task costs of $0.03 for DeepSeek vs $0.07 for Luna, making DeepSeek 2.3x cheaper. @Hesamation added that even with Luna (max) 80% price cut, DeepSeek V4 Flash remains nearly free.
  • Coding and 3D tasks: @cedricchee tested the 248B-parameter DeepSeek-V4-Flash in Codex CLI to generate a 3D voxel pagoda garden, taking 10 minutes and costing only $0.07; another test with 37 API requests took 17 minutes, consuming 3.79M tokens, costing $0.04. @Arindam1729 used it to write a 3D Snake game, completing in 3 iterations for under $1. @Teknium also shared a test of a single complex agent task taking 32 minutes and costing $0.07.
  • Multimodal and security tests: @teortaxesTex found DeepSeek V4-Flash outperformed GPT-5.6 Luna in three Canvas rendering tests at the same price point. However, in a cybersecurity CVE benchmark, DeepSeek Flash v4 retrieved 24/32 CVEs (75% pass@3 recall), still less cost-effective than Luna.
  • Inference speed: @benburtenshaw reported generation speeds up to 400 tokens/s on inference endpoints, with rental costs as low as $10/hour.
  • Local deployment: Multiple developers shared local run experiences. @AmbitiousFold2874 used 4x 5060 Ti (16GB) with DDR4 3200 memory at 128k context; @LegacyRemaster used RTX 6000 (96GB) and W7800 (48GB) hybrid compute for Q3KXL quantized version; @Different-Pickle1021 used A100 (40GB) with Q8KXL quantization (162GB) offloading experts to CPU, achieving 15.8GB VRAM usage; @Nyghtbynger tested at 200K context maintaining coherence; @xjdr noted prompt format and tool-calling quirks may cause integration issues; @zephyrz9 achieved 40 tok/s on M3 Ultra Mac Studio (512GB); @Fit-Produce420 supported 64K context on AMD Strix Halo at 35 tok/s; @akhaliq shared running Q8 quantized version on 7x 3090.

Unconfirmed

  • Some are individual developer tests without full reproduction details; results may vary by task specifics.
  • The intelligence scores (50 vs 51) for DeepSeek V4 Flash and Luna are cited by @zainhas and @joorklee but without specifying evaluation methodology.

Why it matters

DeepSeek V4 Flash offers near-top-tier capabilities at a fraction of the cost, significantly altering the cost structure of AI services. Amid price cuts by OpenAI and others, this model poses substantial competitive pressure on leading vendors and provides developers with a highly cost-effective alternative.

13 more related posts →

Episode 2 · DeepSeek V4-Flash: Agent Leap, Open Source, Ultra-Low Price (2026-07-31, 100 posts)

On July 31, DeepSeek officially released the V4-Flash model API and open-sourced the weights on Hugging Face. The update does not change the architecture but significantly improves agent capabilities through re-post-training, with benchmark scores surpassing the previous V4-Pro-Preview, and offers flagship-level performance at a highly competitive price, sparking community discussion.

Confirmed

  • API and Architecture: The DeepSeek-V4-Flash official API is now in public beta. The model uses a MoE architecture, supports 1M token context, natively supports the OpenAI Responses API format, fully adapts to Codex, and includes a one-click configuration script. The official clarification states that this upgrade is limited to the V4-Flash API; V4-Pro API and App/Web remain unchanged, with V4-Pro official release coming soon.
  • Agent Capability Leap: Officially confirmed, agent capabilities have greatly improved: TerminalBench2.1 score 82.7 (up 25.8 points from the April preview), Cybergym 76.7, Toolathlon 70.3, DeepSWE 54%. It scores 50 on the Artificial Analysis Intelligence Index, up 10 points from the previous generation, surpassing its own V4-Pro by 6 points, and approaching GPT-5.6 Luna and GLM-5.2.
  • Open Source and Ecosystem: DeepSeek released DeepSeek-V4-Flash-0731 weights on Hugging Face under the MIT license, supporting 8-bit and fp8 precision. Hugging Face engineer Victor Mustar deployed a free public inference endpoint that requires no token. Coding tool Cline has already integrated the model.
  • Extreme Cost-Effectiveness: The pricing is highly competitive: $0.11 per million input tokens and $0.27 per million output tokens. According to netizen comparisons, the cost is only 1/50th of Kimi K3 or 1/30th of Gemini 3.6 Flash output price.

Unconfirmed

  • The DeepSWE benchmark result is currently claimed solely by DeepSeek; the poster emphasizes it has not been fully independently verified by third parties.

Why it matters

  • V4-Flash achieves a leap in agent capabilities through re-post-training, providing developers with a more powerful model for task execution and tool calling. Its native Codex compatibility, open-source weights, and highly competitive cost further lower the barrier for agent development, seen by some observers as another "DeepSeek moment."

80 more related posts →

Episode 3 · DeepSeek V4 Series Sparks Debate: Flash Performance Gains and Pro Scale (2026-07-31, 3 posts)

The AI community is actively discussing the DeepSeek V4 series, speculating that V4-Flash's performance gains stem from V4-Pro acting as a reinforcement learning teacher. Debates also focus on whether a 300B parameter V4-Pro can outperform larger 2.8T models.

Episode 4 · DeepSeek V4 Flash Tops OpenRouter, Costs 1% of Rivals (2026-08-03, 14 posts)

DeepSeek V4 Flash (0731) has quickly topped OpenRouter's usage charts, reportedly consuming over 8 trillion tokens daily. Independent benchmarks show it achieves near-top performance at a fraction of the cost of competitors, shifting the focus from raw intelligence to cost-efficiency in long-horizon tasks. This is seen as a direct challenge to global closed-source pricing.

Confirmed

  • Benchmark scores: Per @teortaxesTex relaying WeirdML, DeepSeek v4 Flash (0731) set SOTA records for cost-accuracy with 57.1% and 63.0%, surpassing the recently discounted GPT 5.6 Luna; the evaluator noted scores slightly below expectations but highlighted its agent potential.
  • Cost comparison: Artificial Analysis data (via @SirBoboGargle) shows DeepSeek V4 Flash costs $0.03 on complex real-world workloads vs $3.15 for Claude Fable 5, about 1/100th. @ImaginaryDinner2710 also notes it's 100x cheaper than Anthropic Opus 5.
  • Vals Index: Per @zephyrz9, DeepSeek V4 Flash (0731) is the first model scoring above 60 on Vals at the lowest price, 35x cheaper than the next best, with the advantage almost entirely from coding ability.
  • Usage and migration: @APPSO and @创业邦 report it tops OpenRouter with 8T daily tokens; domestic models' total usage has surpassed US closed-source models for weeks, with many overseas developers migrating production.
  • Concrete pricing: @量子位 reports 7 cents per 3D shooter game, 50 cents per CS game, with overseas resellers adding subsidies, even offering 1-cent-level quotes. @FuSheng0306 estimates heavy users spend 30 yuan/week.
  • Cost-saving practice: @PrajwalTomar notes developers often use expensive flagship models for all tasks, leading to high bills; integrating DeepSeek V4 Flash, which costs 1/50th of flagship models, balances cost and performance.

Why it matters

  • @FuSheng0306 compares the model to the Ford Model T of AI: not necessarily the best performance, but "good enough and affordable for everyone." V4 Flash offers developers a far more economical choice with excellent coding and agent abilities, defining the market's "kill line." When costs drop to a fraction of competitors, the economic model of agent-era long-horizon tasks is redefined—this is no longer just a performance race but a cost-accounting race.

Episode 5 · DeepSeek-V4-Flash Released Across Major Platforms (2026-08-06, 4 posts)

DeepSeek launches the DeepSeek-V4-Flash-0731 model, now available across platforms like Ollama Cloud. With 304B parameters, it outperforms larger models while offering highly competitive API pricing and fast output speeds.

Episode 6 · DeepSeek-V4 Flash Benchmarked: One-Sixth Cost, 80% of Luna's Performance (2026-08-07, 14 posts)

Developer zainhas benchmarked DeepSeek-V4 Flash 0731 against GPT-5.6 Luna on the DeepSWE benchmark for software engineering tasks. The core finding is DeepSeek's exceptional cost-effectiveness: $0.10 per task (about 1/6 of Luna's $0.60) with 80% of Luna's accuracy (53.3% vs 67.2%). Parallel attempts or cascading with Luna can match or exceed Luna while significantly reducing cost. This comparison provides important guidance for model selection, highlighting the potential of low-cost models in software engineering.

Confirmed

  • Cost and accuracy: DeepSeek costs $0.10 per task, Luna $0.60 (5-6x difference); accuracy 53.3% vs 67.2%, Luna only 14 points higher. DeepSeek end-to-end time 23 minutes.
  • Multiple attempts: DeepSeek's economy allows parallel attempts: pass@2 accuracy 81.6% (cost $0.20), pass@4 90.3%; Luna 70.1% and 80.5% respectively.
  • Cascade strategy: DeepSeek first, escalate to Luna on verification failure: accuracy 78.9%, cost $0.385, 37% lower than Luna alone ($0.61), accuracy +11.7%.
  • Task overlap: Both solved 87 tasks; Luna solved 15 unique, DeepSeek only 4; correlation 0.50.
  • Domain performance: Luna wins in 7/8 domains, leading by 30 points in hard areas like program analysis (69-33), concurrency (70-38), runtime internals (86-59); DeepSeek only better in query & config (78-70). DeepSeek collapses on JavaScript (35-60), weak on Python (49-65), decent on Rust (55-60) and Go (62-79).
  • Failure mode: DeepSeek regression breakage rate 9% vs Luna 15%; DeepSeek fails more gracefully, while the pricier model needs regression testing.

Unconfirmed

  • Specific test environment, task set details, and model version info not fully disclosed.

Why it matters

This benchmark gives developers a clear cost-performance trade-off: DeepSeek-V4 Flash is a high-value choice for budget-sensitive scenarios, and via multiple attempts or cascading, it can match or exceed top models without significant cost increase, promoting broader adoption of LLMs in software engineering.