DeepSeek V4 Flash Benchmarks Leak, Outperforming Pro
DeepSeek V4 Flash benchmarks leak, scoring 82.7 on Terminal-Bench 2.1 with a fifth of the parameters of the Pro version, while Unsloth AI's DSpark boosts local inference to 120 tokens/s.
2026-08-05 ~ 2026-08-06 · 3 related posts
- Episode 1: DeepSeek-V4-Flash Architecture Leaked with Million-Token Context(2026-07-31, 3 posts)
- Episode 2: Unsloth Releases Quantized DeepSeek V4 Flash 0731 for Local Deployment(2026-07-31, 6 posts)
- Episode 3: DeepSeek on Huawei Ascend Beats OpenAI in Inference Profitability(2026-08-01, 2 posts)
- Episode 4: DeepSeek-V4-Flash Excels in Frontend Coding with Unmatched Cost-Performance(2026-08-01, 3 posts)
- Episode 5: DeepSeek V4 Preview: Flash to Introduce Four-Level Reasoning Effort(2026-08-01, 2 posts)
- Episode 6: DeepSeek Drastically Reduces Training Compute Costs Across Models(2026-08-01, 2 posts)
- Episode 7: DeepSeek V4-Flash API Public Beta Launches with Major Agent Upgrades(2026-08-01, 5 posts)
- Episode 8: DeepSeek V4-Flash Costs 105x Less, But Stability and Benchmark Overfitting Questioned(2026-08-01, 17 posts)
- Episode 9: DeepSeek V4 Flash Sparks a Wave of Local Deployment Tests(2026-08-01, 26 posts)
- Episode 10: DeepSeek V4 Flash Quantization Tests Show Major Speed Gains Without Quality Loss(2026-08-04, 2 posts)
- Episode 11: Single RTX 5090 Runs DeepSeek 1M Context(2026-08-04, 2 posts)
- Episode 12: DeepSeek-V4-Flash Tops Cost-Efficiency with Ultra-Low Running Costs(2026-08-05, 4 posts)
- Episode 13: DeepSeek V4 Flash Benchmarks Leak, Outperforming Pro(2026-08-05, 3 posts)
- Episode 14: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance, Costing a Quarter of GPT-5.6(2026-08-08, 9 posts)
- Episode 15: DeepSeek V4 Flash Passes 22 Coding Tests on Dual DGX Spark Cluster(2026-08-10, 2 posts)
- DeepSeek V4 Flash Local Benchmark: MXFP4 Quantization Balances Speed and Top Scores — WonderRico · 2026-08-05
- DeepSeek V4 Flash Benchmarks Leak: Beats Pro with 1/5 Parameters — togethercompute · 2026-08-06
- DeepSeek-V4-Flash Local Inference Hits 120 tokens/s via Unsloth — danielhanchen · 2026-08-06