FULL STORY
DeepSeek V4 Flash: Open-Source Disruption and Tests
DeepSeek V4 Flash sparked a wave of testing and quickly topped platform usage charts, offering top-tier performance at a fraction of competitors' costs.
2026-07-31 ~ 2026-08-08 · 6 episodes · 168 posts
Episode 1 · DeepSeek V4 Flash Hands-on: Low Cost, High Performance Sparks Community Buzz (2026-07-31, 33 posts)
DeepSeek-V4-Flash-0731 has sparked extensive community testing due to its ultra-low cost and near-top-tier performance. Benchmarks show a 50 intelligence score, matching Gemini 3.6 Flash, with real-world capabilities close to Claude 3.5 Sonnet. In multiple complex task tests, it demonstrated exceptional cost-effectiveness in the AI price war.
Confirmed
- Performance benchmarks: According to Artificial Analysis, DeepSeek V4 Flash (0731) scored 50, matching Gemini 3.6 Flash and approaching the record 51 set in March 2026. @bindureddy noted its real-world capability is close to Claude 3.5 Sonnet.
- Cost advantage: @teortaxesTex pointed out that while GPT-5.6 Luna has low input token prices, its output token price ($1.20/M) is much higher than DeepSeek's ($0.28/M). @autooff estimated that for 7 billion tokens per month, Luna would cost $250, while DeepSeek would cost only a third. @zainhas's comparison showed per-task costs of $0.03 for DeepSeek vs $0.07 for Luna, making DeepSeek 2.3x cheaper. @Hesamation added that even with Luna (max) 80% price cut, DeepSeek V4 Flash remains nearly free.
- Coding and 3D tasks: @cedricchee tested the 248B-parameter DeepSeek-V4-Flash in Codex CLI to generate a 3D voxel pagoda garden, taking 10 minutes and costing only $0.07; another test with 37 API requests took 17 minutes, consuming 3.79M tokens, costing $0.04. @Arindam1729 used it to write a 3D Snake game, completing in 3 iterations for under $1. @Teknium also shared a test of a single complex agent task taking 32 minutes and costing $0.07.
- Multimodal and security tests: @teortaxesTex found DeepSeek V4-Flash outperformed GPT-5.6 Luna in three Canvas rendering tests at the same price point. However, in a cybersecurity CVE benchmark, DeepSeek Flash v4 retrieved 24/32 CVEs (75% pass@3 recall), still less cost-effective than Luna.
- Inference speed: @benburtenshaw reported generation speeds up to 400 tokens/s on inference endpoints, with rental costs as low as $10/hour.
- Local deployment: Multiple developers shared local run experiences. @AmbitiousFold2874 used 4x 5060 Ti (16GB) with DDR4 3200 memory at 128k context; @LegacyRemaster used RTX 6000 (96GB) and W7800 (48GB) hybrid compute for Q3KXL quantized version; @Different-Pickle1021 used A100 (40GB) with Q8KXL quantization (162GB) offloading experts to CPU, achieving 15.8GB VRAM usage; @Nyghtbynger tested at 200K context maintaining coherence; @xjdr noted prompt format and tool-calling quirks may cause integration issues; @zephyrz9 achieved 40 tok/s on M3 Ultra Mac Studio (512GB); @Fit-Produce420 supported 64K context on AMD Strix Halo at 35 tok/s; @akhaliq shared running Q8 quantized version on 7x 3090.
Unconfirmed
- Some are individual developer tests without full reproduction details; results may vary by task specifics.
- The intelligence scores (50 vs 51) for DeepSeek V4 Flash and Luna are cited by @zainhas and @joorklee but without specifying evaluation methodology.
Why it matters
DeepSeek V4 Flash offers near-top-tier capabilities at a fraction of the cost, significantly altering the cost structure of AI services. Amid price cuts by OpenAI and others, this model poses substantial competitive pressure on leading vendors and provides developers with a highly cost-effective alternative.
- DeepSeek Flash Offers Dirt-Cheap Pricing and Solid Performance — bindureddy · 2026-07-31
- Luna Offers Cheaper Inputs, but DeepSeek Wins Cache Economics in Long Agentic Sessions — teortaxesTex · 2026-07-31
- DeepSeek-V4-Flash Ported to Run on AMD Strix Halo APU — Fit-Produce420 · 2026-07-31
- Integrating DeepSeek-V4-Flash into Codex: Costs 89x Less Than Opus — teortaxesTex · 2026-07-31
- DeepSeek V4-Flash Beats GPT-5.6 Luna in Multimodal Canvas Tests at Same Price — teortaxesTex · 2026-07-31
- DeepSeek-V4-Flash Agent Eval: Completes 3D Task for $0.07 — cedric_chee · 2026-07-31
- DeepSeek-V4-Flash Tested in Codex CLI: 3D Task Costs Only $0.07 — cedric_chee · 2026-07-31
- DeepSeek Costs 1/3 of GPT Luna for Coding: A Practical Token & Expense Breakdown — auto_off · 2026-07-31
- Testing DeepSeek-V4-Flash: 37 API Calls in Codex CLI Cost Only $0.04 — cedric_chee · 2026-07-31
- DeepSeek V4 Flash Coding Performance Falls Between Opus 4.8 and 5 — corruptbytes · 2026-07-31
- DeepSeek-V4-Flash Runs at 400 tps for Just $10/hour on Inference Endpoints — ben_burtenshaw · 2026-07-31
- Testing DeepSeek V4 Flash on a 3D Game: 3 Iterations, Under $1 Cost — Arindam_1729 · 2026-07-31
- 7x RTX 3090s Barely Run DeepSeek-V4 Q8 Quantization Locally — _akhaliq · 2026-08-01
- DeepSeek v4-flash tested: strong performance but integration quirks remain — _xjdr · 2026-08-01
- DeepSeek New Model Tested: Great Long-Context, But Tool Calling Quirks — TheZachMueller · 2026-08-01
- Hands-on: DeepSeek V4 Flash Stays Coherent at 200K Context, Excels in Reasoning — Nyghtbynger · 2026-08-01
- DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s — Different-Pickle1021 · 2026-08-01
- DeepSeek Flash v4 Cybersecurity Test: Finds 24 CVEs but Loses on Cost-Efficiency to Luna — teortaxesTex · 2026-08-01
- DeepSeek V4 Flash undercuts GPT-5.6 Luna: 2.3x cheaper with similar intelligence — zainhas · 2026-08-01
- DeepSeek V4 Flash is Basically Free, Luna Max Offers Insane Value After 80% Price Cut — Hesamation · 2026-08-01
Episode 2 · DeepSeek V4-Flash: Agent Leap, Open Source, Ultra-Low Price (2026-07-31, 100 posts)
On July 31, DeepSeek officially released the V4-Flash model API and open-sourced the weights on Hugging Face. The update does not change the architecture but significantly improves agent capabilities through re-post-training, with benchmark scores surpassing the previous V4-Pro-Preview, and offers flagship-level performance at a highly competitive price, sparking community discussion.
Confirmed
- API and Architecture: The DeepSeek-V4-Flash official API is now in public beta. The model uses a MoE architecture, supports 1M token context, natively supports the OpenAI Responses API format, fully adapts to Codex, and includes a one-click configuration script. The official clarification states that this upgrade is limited to the V4-Flash API; V4-Pro API and App/Web remain unchanged, with V4-Pro official release coming soon.
- Agent Capability Leap: Officially confirmed, agent capabilities have greatly improved: TerminalBench2.1 score 82.7 (up 25.8 points from the April preview), Cybergym 76.7, Toolathlon 70.3, DeepSWE 54%. It scores 50 on the Artificial Analysis Intelligence Index, up 10 points from the previous generation, surpassing its own V4-Pro by 6 points, and approaching GPT-5.6 Luna and GLM-5.2.
- Open Source and Ecosystem: DeepSeek released DeepSeek-V4-Flash-0731 weights on Hugging Face under the MIT license, supporting 8-bit and fp8 precision. Hugging Face engineer Victor Mustar deployed a free public inference endpoint that requires no token. Coding tool Cline has already integrated the model.
- Extreme Cost-Effectiveness: The pricing is highly competitive: $0.11 per million input tokens and $0.27 per million output tokens. According to netizen comparisons, the cost is only 1/50th of Kimi K3 or 1/30th of Gemini 3.6 Flash output price.
Unconfirmed
- The DeepSWE benchmark result is currently claimed solely by DeepSeek; the poster emphasizes it has not been fully independently verified by third parties.
Why it matters
- V4-Flash achieves a leap in agent capabilities through re-post-training, providing developers with a more powerful model for task execution and tool calling. Its native Codex compatibility, open-source weights, and highly competitive cost further lower the barrier for agent development, seen by some observers as another "DeepSeek moment."
- DeepSeek-V4-Flash Officially Released, Surpassing Pro Preview in Agent Benchmarks — 赛博禅心 · 2026-07-31
- DeepSeek updates V4-Flash, teases imminent V4-Pro release — Nunki08 · 2026-07-31
- DeepSeek-V4-Flash Official API Released with Major Agent Upgrades — APPSO · 2026-07-31
- DeepSeek-V4-Flash Updated with Significant Score Improvements — cedric_chee · 2026-07-31
- DeepSeek V4 Flash Model Goes Live on Official API — ElMess-Siah · 2026-07-31
- DeepSeek V4 Flash Official Release Outperforms V4 Pro Preview in Agent Capabilities — 歸藏的AI工具箱 · 2026-07-31
- DeepSeek V4 Flash Officially Released with Major Agent Capability Boost — op7418 · 2026-07-31
- Updated DeepSeek V4 Flash Scores 54% on DeepSWE Benchmark — zainhas · 2026-07-31
- DeepSeek-V4-Flash Official API Released with Major Agent Upgrades — zainhas · 2026-07-31
- DeepSeek V4 Flash official release: big agent gains, cost advantage — 数字生命卡兹克 · 2026-07-31
- DeepSeek-V4-Flash Officially Released with Major Agent Upgrades — jiqizhixin · 2026-07-31
- DeepSeek Clarifies V4-Flash Upgrade Details, V4-Pro Coming Soon — deepseek_ai · 2026-07-31
- DeepSeek V4-Flash Natively Supports Codex with Setup Guide — zephyr_z9 · 2026-07-31
- DeepSeek Quietly Launches V4-Flash API; Open Weights Expected — Hot_Example_4456 · 2026-07-31
- DeepSeek-V4-Flash Tested: Blazing Fast Speed and Outperforms Pro — vista8 · 2026-07-31
- DeepSeek V4-Flash Beats Pro in Agent Tests, Offers 1-Click Codex Integration — 卡尔的AI沃茨 · 2026-07-31
- DeepSeek Releases New DeepSeek-V4-Flash-0731 Model — deepseek-ai · 2026-07-31
- DeepSeek V4-Flash Scores 50 on Artificial Analysis Index, 1 Point Below GLM-5.2 — MagicZhang · 2026-07-31
- DeepSeek-V4-Flash Official API Launches with Enhanced Agent Capabilities and Ultra-low Pricing — 智东西 · 2026-07-31
- DeepSeek V4-Flash Hits Index Score of 50 at a Cost of Just $0.20 — xeophon · 2026-07-31
Episode 3 · DeepSeek V4 Series Sparks Debate: Flash Performance Gains and Pro Scale (2026-07-31, 3 posts)
The AI community is actively discussing the DeepSeek V4 series, speculating that V4-Flash's performance gains stem from V4-Pro acting as a reinforcement learning teacher. Debates also focus on whether a 300B parameter V4-Pro can outperform larger 2.8T models.
- DeepSeek V4-Flash Performance Jump May Stem from V4-Pro as RL Teacher — teortaxesTex · 2026-07-31
- Debate: Can a 300B V4 Pro Model Beat a Newly Released 2.8T Model? — scaling01 · 2026-08-01
- V4-Pro Improvement Debate: Stronger Base vs. Saturated Evals — teortaxesTex · 2026-08-01
Episode 4 · DeepSeek V4 Flash Tops OpenRouter, Costs 1% of Rivals (2026-08-03, 14 posts)
DeepSeek V4 Flash (0731) has quickly topped OpenRouter's usage charts, reportedly consuming over 8 trillion tokens daily. Independent benchmarks show it achieves near-top performance at a fraction of the cost of competitors, shifting the focus from raw intelligence to cost-efficiency in long-horizon tasks. This is seen as a direct challenge to global closed-source pricing.
Confirmed
- Benchmark scores: Per @teortaxesTex relaying WeirdML, DeepSeek v4 Flash (0731) set SOTA records for cost-accuracy with 57.1% and 63.0%, surpassing the recently discounted GPT 5.6 Luna; the evaluator noted scores slightly below expectations but highlighted its agent potential.
- Cost comparison: Artificial Analysis data (via @SirBoboGargle) shows DeepSeek V4 Flash costs $0.03 on complex real-world workloads vs $3.15 for Claude Fable 5, about 1/100th. @ImaginaryDinner2710 also notes it's 100x cheaper than Anthropic Opus 5.
- Vals Index: Per @zephyrz9, DeepSeek V4 Flash (0731) is the first model scoring above 60 on Vals at the lowest price, 35x cheaper than the next best, with the advantage almost entirely from coding ability.
- Usage and migration: @APPSO and @创业邦 report it tops OpenRouter with 8T daily tokens; domestic models' total usage has surpassed US closed-source models for weeks, with many overseas developers migrating production.
- Concrete pricing: @量子位 reports 7 cents per 3D shooter game, 50 cents per CS game, with overseas resellers adding subsidies, even offering 1-cent-level quotes. @FuSheng0306 estimates heavy users spend 30 yuan/week.
- Cost-saving practice: @PrajwalTomar notes developers often use expensive flagship models for all tasks, leading to high bills; integrating DeepSeek V4 Flash, which costs 1/50th of flagship models, balances cost and performance.
Why it matters
- @FuSheng0306 compares the model to the Ford Model T of AI: not necessarily the best performance, but "good enough and affordable for everyone." V4 Flash offers developers a far more economical choice with excellent coding and agent abilities, defining the market's "kill line." When costs drop to a fraction of competitors, the economic model of agent-era long-horizon tasks is redefined—this is no longer just a performance race but a cost-accounting race.
- DeepSeek V4 Flash Burns 8T Tokens Daily, Reshaping Agent-Era Model Pricing — APPSO · 2026-08-03
- DeepSeek-V4-Flash tops OpenRouter, reshaping LLM pricing with extreme cost-efficiency — 创业邦 · 2026-08-03
- DeepSeek v4 Flash Sets New Cost/Accuracy SOTA on WeirdML — teortaxesTex · 2026-08-03
- DeepSeek V4 Flash tops the Vals Index above 60 at 35x lower cost — zephyr_z9 · 2026-08-04
- DeepSeek V4-Flash is being used for 3D games at cents-level costs — 量子位 · 2026-08-04
- DeepSeek V4 Costs 1% of Claude: China's AI Price War Disrupts the Market — SirBoboGargle · 2026-08-04
- Stop Burning Money on Opus: The DeepSeek V4 Flash Fix — PrajwalTomar_ · 2026-08-04
- DeepSeek Dubbed the 'Model T' of the AI Era for Its Aggressive Pricing — FuSheng_0306 · 2026-08-04
- DeepSeek V4 Flash Costs 1/100th of Top Models, Set to Trigger Autonomous Agent Wave — Imaginary_Dinner2710 · 2026-08-05
- DeepSeek Flash's performance is mind-boggling,网友直呼“不合理” — yacineMTB · 2026-08-05
- Developer Marvels at DeepSeek's Pricing: Indefinite Use for Almost Zero Cost — yacineMTB · 2026-08-05
- DeepSeek V4 Flash Matches GLM 5.2, Tops OpenRouter Rankings — natolambert · 2026-08-05
- DeepSeek-V4-Flash Called a 'Monster' Model, Tops OpenRouter Usage — natolambert · 2026-08-05
- Developer Predicts DeepSeek V4 Flash Will Dominate Claude Due to 500x Cost Efficiency — tekbog · 2026-08-05
Episode 5 · DeepSeek-V4-Flash Released Across Major Platforms (2026-08-06, 4 posts)
DeepSeek launches the DeepSeek-V4-Flash-0731 model, now available across platforms like Ollama Cloud. With 304B parameters, it outperforms larger models while offering highly competitive API pricing and fast output speeds.
- DeepSeek-V4-Flash Launches: 304B Params Outperforms Larger Models at $0.14 — JeremyCMorgan · 2026-08-06
- DeepSeek v4 Flash API Pricing Revealed: Nearly 30% Cheaper — youjiaxuan · 2026-08-07
- DeepSeek V4 Flash launches on engy.ai, 68% cheaper than official API — markjeffrey · 2026-08-08
- DeepSeek-V4-Flash Rolls Out on Ollama Cloud with 120+ Output TPS — ollama · 2026-08-08
Episode 6 · DeepSeek-V4 Flash Benchmarked: One-Sixth Cost, 80% of Luna's Performance (2026-08-07, 14 posts)
Developer zainhas benchmarked DeepSeek-V4 Flash 0731 against GPT-5.6 Luna on the DeepSWE benchmark for software engineering tasks. The core finding is DeepSeek's exceptional cost-effectiveness: $0.10 per task (about 1/6 of Luna's $0.60) with 80% of Luna's accuracy (53.3% vs 67.2%). Parallel attempts or cascading with Luna can match or exceed Luna while significantly reducing cost. This comparison provides important guidance for model selection, highlighting the potential of low-cost models in software engineering.
Confirmed
- Cost and accuracy: DeepSeek costs $0.10 per task, Luna $0.60 (5-6x difference); accuracy 53.3% vs 67.2%, Luna only 14 points higher. DeepSeek end-to-end time 23 minutes.
- Multiple attempts: DeepSeek's economy allows parallel attempts: pass@2 accuracy 81.6% (cost $0.20), pass@4 90.3%; Luna 70.1% and 80.5% respectively.
- Cascade strategy: DeepSeek first, escalate to Luna on verification failure: accuracy 78.9%, cost $0.385, 37% lower than Luna alone ($0.61), accuracy +11.7%.
- Task overlap: Both solved 87 tasks; Luna solved 15 unique, DeepSeek only 4; correlation 0.50.
- Domain performance: Luna wins in 7/8 domains, leading by 30 points in hard areas like program analysis (69-33), concurrency (70-38), runtime internals (86-59); DeepSeek only better in query & config (78-70). DeepSeek collapses on JavaScript (35-60), weak on Python (49-65), decent on Rust (55-60) and Go (62-79).
- Failure mode: DeepSeek regression breakage rate 9% vs Luna 15%; DeepSeek fails more gracefully, while the pricier model needs regression testing.
Unconfirmed
- Specific test environment, task set details, and model version info not fully disclosed.
Why it matters
This benchmark gives developers a clear cost-performance trade-off: DeepSeek-V4 Flash is a high-value choice for budget-sensitive scenarios, and via multiple attempts or cascading, it can match or exceed top models without significant cost increase, promoting broader adoption of LLMs in software engineering.
- DeepSeek-V4 vs GPT-5.6: 1/6 the Cost, 80% the Quality on Coding Tasks — zainhas · 2026-08-07
- Luna Costs 6x More for Only 14pts Accuracy Gain; DeepSeek Wins on Value — zainhas · 2026-08-07
- DeepSeek at $0.10/Task: 5x Cheaper Than Luna — zainhas · 2026-08-07
- DeepSeek's Parallel Attempts Catch Up: pass@2 Beats Luna pass@1 at $0.20 — zainhas · 2026-08-07
- DeepSeek Fails More Gracefully: 9% Regression Rate vs Luna's 15% — zainhas · 2026-08-07
- Luna Wins 7/8 Domains, DeepSeek Only Holds Query & Config — zainhas · 2026-08-07
- DeepSeek's JavaScript Performance Collapses, Rust and Go Fine — zainhas · 2026-08-07
- DeepSeek vs Luna: High Overlap, Luna Solves 15 Unique Tasks, DeepSeek Only 4 — zainhas · 2026-08-07
- DeepSeek+Luna Cascade: 78.9% Accuracy at 63% Cost, Beats Luna Alone — zainhas · 2026-08-07
- DeepSeek + Luna Combo Strategy: 37% Cheaper, 11.7% More Accurate — zainhas · 2026-08-07
- DeepSeek-V4 Flash Delivers 80% of GPT-5.6 Luna Performance at 1/6 Cost — zainhas · 2026-08-07
- DeepSeek Flash Hits 80% of GPT 5.6 Luna Quality at 1/6 the Cost — pbaylies · 2026-08-07
- DeepSeek V4 Beats GPT 5.6 in Coding Cost-Efficiency by 6x — togethercompute · 2026-08-07
- DeepSeek Cascade Beats GPT-5.6 Luna on DeepSWE at 37% Lower Cost — togethercompute · 2026-08-08