DeepSeek V4 Flash 0731 Scores 50 on Intelligence Index, Closing Gap with GPT-5.6
ArtificialAnlys · x · 2026-07-31
Artificial Analysis released its latest evaluation of DeepSeek V4 Flash 0731. The model scores 50 on the Intelligence Index, a 10-point jump that puts it 6 points ahead of DeepSeek V4 Pro, landing perfectly on the Pareto frontier for intelligence vs. cost.
Key Data & Highlights:
- Extreme Cost-Effectiveness: Thanks to a 98% cache hit discount, its cost per task is about 60% lower than the similarly intelligent GPT-5.6 Luna.
- Massive Agentic Gains: Elo rating on GDPval-AA v2 surged from 1189 to 1559, with significant improvements in Terminal-Bench and other tests.
- Token Efficiency: Output tokens used for the Intelligence Index dropped 12% compared to its predecessor.
- Reduced Hallucinations: The AA-Omniscience Hallucination Rate fell by 12 points to 84%, driving overall improvement while maintaining a 284B parameter size.
The model retains a 1M token context window, with full weights expected to be released in the coming weeks.
Related event: DeepSeek V4-Flash Benchmarks Surge, Approaching GPT-5.6(18 posts)→
More from Models
- DeepSeek V4-Flash Beats GPT-5.6 Luna in Multimodal Canvas Tests at Same Price — teortaxesTex · 2026-07-31
- DeepSeek V4-Flash Performance Jump May Stem from V4-Pro as RL Teacher — teortaxesTex · 2026-07-31
- DeepSeek Praised for World-Class RL Training That Avoids Hallucinations — teortaxesTex · 2026-07-31
- Hands-on: OpenAI's o3 Remains a Beast for OSINT Tasks — bytebot · 2026-07-31
- Rumor: Gemini 3.5 Pro Could Drop This Weekend — gaganghotra_ · 2026-07-31
- Kimi's RL Training Praised: Effectively Avoids Hallucinations and Verbosity — teortaxesTex · 2026-07-31