DeepSeek-V4-Flash-0731 Benchmarks: Matches Top Closed Models at Fraction of Cost
Following the release of DeepSeek V4 Flash, multiple developers conducted hands-on tests. Results show that the model performs impressively in tasks like coding, 3D generation, and cybersecurity, at a fraction of the cost of OpenAI's GPT-5.6 Luna. These tests highlight DeepSeek's significant advantage in cost-effectiveness, sparking community discussions about the competitiveness of open-source models.
Confirmed
- Cost Comparison: @teortaxesTex pointed out that Luna's output token price ($1.20/M) is significantly higher than DeepSeek's ($0.28/M). @autooff estimated that for a monthly usage of 7 billion tokens, Luna would cost about $250, whereas DeepSeek is only about a third of that. Data from @zainhas shows the cost per task is $0.03 for DeepSeek V4 Flash and $0.07 for Luna, making it 2.3 times cheaper.
- Coding and 3D Tasks: @cedricchee tested DeepSeek-V4-Flash (248B parameters) in Codex CLI to generate a 3D voxel pagoda garden for just $0.07; @Arindam1729 used it to code a 3D Snake game, completing it in 3 iterations for a total cost of under $1. In another test by @cedricchee, 37 API requests cost only $0.04.
- Cybersecurity: In the CVE benchmark, @teortaxesTex noted that DeepSeek Flash v4 recovered 24/32 CVEs with a pass@3 recall rate of 75%, though its cost-effectiveness still fell slightly short of Luna.
- Inference Speed: @benburtenshaw reported that DeepSeek-V4-Flash-0731 can reach 400 tokens/s on inference endpoints, with a rental cost of just $10/hour.
- Multimodal Comparison: In a price-matched test by @teortaxesTex, DeepSeek V4-Flash outperformed GPT-5.6 Luna in a Canvas rendering test.
Unconfirmed
- Some tests are personal benchmarks by developers without full reproduction details, so results may vary depending on the specific task.
- The intelligence scores comparing DeepSeek V4 Flash and Luna (50 vs 51) came from @zainhas, but the specific evaluation methodology was not disclosed.
Why It Matters
Offering capabilities that closely rival or even exceed closed-source models at an extremely low price, DeepSeek V4 Flash could disrupt the cost structure of AI services and put competitive pressure on players like OpenAI. The wealth of hands-on testing data from the developer community also provides a solid reference for model selection.
2026-07-31 ~ 2026-08-01 · 31 related posts
Primary sources
- DeepSeek Flash Offers Dirt-Cheap Pricing and Solid Performance — bindureddy · 2026-07-31
- [source] Luna Offers Cheaper Inputs, but DeepSeek Wins Cache Economics in Long Agentic Sessions — teortaxesTex · 2026-07-31
- DeepSeek-V4-Flash Ported to Run on AMD Strix Halo APU — Fit-Produce420 · 2026-07-31
- Integrating DeepSeek-V4-Flash into Codex: Costs 89x Less Than Opus — teortaxesTex · 2026-07-31
- DeepSeek V4-Flash Beats GPT-5.6 Luna in Multimodal Canvas Tests at Same Price — teortaxesTex · 2026-07-31
- [source] DeepSeek-V4-Flash Agent Eval: Completes 3D Task for $0.07 — cedric_chee · 2026-07-31
- DeepSeek-V4-Flash Tested in Codex CLI: 3D Task Costs Only $0.07 — cedric_chee · 2026-07-31
- DeepSeek Costs 1/3 of GPT Luna for Coding: A Practical Token & Expense Breakdown — auto_off · 2026-07-31
- Testing DeepSeek-V4-Flash: 37 API Calls in Codex CLI Cost Only $0.04 — cedric_chee · 2026-07-31
- DeepSeek V4 Flash Coding Performance Falls Between Opus 4.8 and 5 — corruptbytes · 2026-07-31
- DeepSeek-V4-Flash Runs at 400 tps for Just $10/hour on Inference Endpoints — ben_burtenshaw · 2026-07-31
- Testing DeepSeek V4 Flash on a 3D Game: 3 Iterations, Under $1 Cost — Arindam_1729 · 2026-07-31
- 7x RTX 3090s Barely Run DeepSeek-V4 Q8 Quantization Locally — _akhaliq · 2026-08-01
- DeepSeek v4-flash tested: strong performance but integration quirks remain — _xjdr · 2026-08-01
- DeepSeek New Model Tested: Great Long-Context, But Tool Calling Quirks — TheZachMueller · 2026-08-01
- Hands-on: DeepSeek V4 Flash Stays Coherent at 200K Context, Excels in Reasoning — Nyghtbynger · 2026-08-01
- DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s — Different-Pickle1021 · 2026-08-01
- DeepSeek Flash v4 Cybersecurity Test: Finds 24 CVEs but Loses on Cost-Efficiency to Luna — teortaxesTex · 2026-08-01
- DeepSeek V4 Flash undercuts GPT-5.6 Luna: 2.3x cheaper with similar intelligence — zainhas · 2026-08-01
- DeepSeek V4 Flash is Basically Free, Luna Max Offers Insane Value After 80% Price Cut — Hesamation · 2026-08-01
- Optimizing DeepSeek V4 Flash on 4x 5060 Ti: A Local Deployment Discussion — Ambitious_Fold_2874 · 2026-08-01
- DeepSeek V4 Flash Scores 50 on Intelligence Index, Undercutting GPT-5.6 by 60% — solyarisoftware · 2026-08-01
- DeepSeek-V4-Flash tested: hits 2,100 tok/s aggregate on 4x RTX PRO 6000 — BanghuaZ · 2026-08-01
- DeepSeek Flash First Look: Rivals Sonnet at 1/6 the Cost — bindureddy · 2026-08-01
- DeepSeek Hits 40 tok/s Locally on M3 Ultra Mac Studio — zephyr_z9 · 2026-08-01
- Benchmarks: Running DeepSeek Locally on 4x 5060 Ti with 128k Context — Ambitious_Fold_2874 · 2026-08-01
- Teknium Tests DeepSeek V4 Flash: Full Agent Task Costs Just $0.07 — Teknium · 2026-08-01
- [source] DeepSeek V4 Flash local benchmark nearly matches top frontier models from 5 months ago — joorklee · 2026-08-01
- DeepSeek V4 Flash 0731 vs. ChatGPT Luna: Hands-on Comparison — perelmanych · 2026-08-01
- Can a Single RTX PRO 6000 Run DeepSeek V4 Flash Locally? — TechNerd10191 · 2026-08-01
- DeepSeek-V4-Flash Runs Complex Physics Sim Locally with Imppressive Results — LegacyRemaster · 2026-08-01