DeepSeek V4 Flash Benchmarks: Low Cost, High Performance, Outshines Luna
Following the release of DeepSeek V4 Flash, multiple developers conducted hands-on tests. Results show that the model performs impressively in tasks like coding, 3D generation, and cybersecurity, at a fraction of the cost of OpenAI's GPT-5.6 Luna. These tests highlight DeepSeek's significant advantage in cost-effectiveness, sparking community discussions about the competitiveness of open-source models.
Confirmed
- Cost Comparison: @teortaxesTex pointed out that Luna's output token price ($1.20/M) is significantly higher than DeepSeek's ($0.28/M). @autooff estimated that for a monthly usage of 7 billion tokens, Luna would cost about $250, whereas DeepSeek is only about a third of that. Data from @zainhas shows the cost per task is $0.03 for DeepSeek V4 Flash and $0.07 for Luna, making it 2.3 times cheaper.
- Coding and 3D Tasks: @cedricchee tested DeepSeek-V4-Flash (248B parameters) in Codex CLI to generate a 3D voxel pagoda garden for just $0.07; @Arindam1729 used it to code a 3D Snake game, completing it in 3 iterations for a total cost of under $1. In another test by @cedricchee, 37 API requests cost only $0.04.
- Cybersecurity: In the CVE benchmark, @teortaxesTex noted that DeepSeek Flash v4 recovered 24/32 CVEs with a pass@3 recall rate of 75%, though its cost-effectiveness still fell slightly short of Luna.
- Inference Speed: @benburtenshaw reported that DeepSeek-V4-Flash-0731 can reach 400 tokens/s on inference endpoints, with a rental cost of just $10/hour.
- Multimodal Comparison: In a price-matched test by @teortaxesTex, DeepSeek V4-Flash outperformed GPT-5.6 Luna in a Canvas rendering test.
Unconfirmed
- Some tests are personal benchmarks by developers without full reproduction details, so results may vary depending on the specific task.
- The intelligence scores comparing DeepSeek V4 Flash and Luna (50 vs 51) came from @zainhas, but the specific evaluation methodology was not disclosed.
Why It Matters
Offering capabilities that closely rival or even exceed closed-source models at an extremely low price, DeepSeek V4 Flash could disrupt the cost structure of AI services and put competitive pressure on players like OpenAI. The wealth of hands-on testing data from the developer community also provides a solid reference for model selection.
2026-07-31 ~ 2026-08-01 · 12 related posts
Primary sources
- Luna Offers Cheaper Inputs, but DeepSeek Wins Cache Economics in Long Agentic Sessions — teortaxesTex ·
- DeepSeek Costs 1/3 of GPT Luna for Coding: A Practical Token & Expense Breakdown — auto_off ·
- DeepSeek Flash v4 Cybersecurity Test: Finds 24 CVEs but Loses on Cost-Efficiency to Luna — teortaxesTex ·
- DeepSeek Flash Offers Dirt-Cheap Pricing and Solid Performance — bindureddy · 2026-07-31
- [source] Luna Offers Cheaper Inputs, but DeepSeek Wins Cache Economics in Long Agentic Sessions — teortaxesTex · 2026-07-31
- Integrating DeepSeek-V4-Flash into Codex: Costs 89x Less Than Opus — teortaxesTex · 2026-07-31
- DeepSeek V4-Flash Beats GPT-5.6 Luna in Multimodal Canvas Tests at Same Price — teortaxesTex · 2026-07-31
- DeepSeek-V4-Flash Agent Eval: Completes 3D Task for $0.07 — cedric_chee · 2026-07-31
- DeepSeek-V4-Flash Tested in Codex CLI: 3D Task Costs Only $0.07 — cedric_chee · 2026-07-31
- [source] DeepSeek Costs 1/3 of GPT Luna for Coding: A Practical Token & Expense Breakdown — auto_off · 2026-07-31
- Testing DeepSeek-V4-Flash: 37 API Calls in Codex CLI Cost Only $0.04 — cedric_chee · 2026-07-31
- DeepSeek-V4-Flash Runs at 400 tps for Just $10/hour on Inference Endpoints — ben_burtenshaw · 2026-07-31
- Testing DeepSeek V4 Flash on a 3D Game: 3 Iterations, Under $1 Cost — Arindam_1729 · 2026-07-31
- [source] DeepSeek Flash v4 Cybersecurity Test: Finds 24 CVEs but Loses on Cost-Efficiency to Luna — teortaxesTex · 2026-08-01
- DeepSeek V4 Flash undercuts GPT-5.6 Luna: 2.3x cheaper with similar intelligence — zainhas · 2026-08-01