DeepSeek V4 Flash scores 89% on ARC-AGI-1 benchmark
teortaxesTex · x · 2026-08-09
DeepSeek V4 Flash 0731 demonstrates strong performance on the ARC-AGI benchmark. At max effort, the model scores 89.0% on the ARC-AGI-1 Semi-Private set ($0.02 per task) and 61.4% on ARC-AGI-2 Semi-Private ($0.04 per task). The original poster noted the irony of AI evaluation timelines, remarking that by the time the highly intensive ARC-AGI-3 evaluation is completed, the model being tested will likely already be obsolete.
More from Models
- Bindu Reddy's Top Tier Model List Ranks Fable 5, Sol, and Opus 4.8 as S-Tier — bindureddy · 2026-08-26
- Gemini glitch: Model ignores context completely — WishIWasALemon · 2026-08-26
- Ornith-1.5 Open Models Released, Claiming Claude Opus Performance — alejandroll10 · 2026-08-26
- 14-year AI veteran: Grok understood code I thought no one ever would — Kuprel · 2026-08-26
- Together Ranks Top Open Models: Kimi K3 and DeepSeek V4 Lead Use Cases — togethercompute · 2026-08-26
- Questions over Astra's progress: 2 months for 3 more models? — teortaxesTex · 2026-08-26