DeepSeek V4 Flash Dominates ARC-AGI Cost-Performance Frontier at 1/4 the Cost
GregKamradt · x · 2026-08-08
DeepSeek V4 Flash has delivered an outstanding performance on the ARC-AGI benchmark, setting a new standard on the cost-to-performance Pareto frontier.
According to verified results from ARC Prize, the model scored 61.4% on ARC-AGI-2 ($0.04/task) and 89.0% on ARC-AGI-1 ($0.02/task). This effectively matches the performance of GPT-5.6 Luna (Max) at just a quarter of the cost.
Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance Frontier(2 posts)→
More from Models
- Qwen3.8-Max beats Gemini 3.5 Flash by 8.5 points in tests — usamawahabkhan · 2026-08-08
- Semianalysis Deep Dive on Gemini 3.5 Pro: Performance and Architecture Insights — Charuru · 2026-08-08
- DeepSeek V4 Flash Appears on ARC Prize Leaderboard — tosh · 2026-08-08
- Pokee AI Launches Isaac Model with 10M-Token Context and API — Kyrannio · 2026-08-08
- Opinion: Google or Meta Could Win the AI Race by Dropping a Better Open-Source Model Than K3 — bindureddy · 2026-08-08
- Why LLMs Can't Count Tokens: The Need for Intermediate Steps — ctjlewis · 2026-08-08