DeepSeek V4 Flash scores 89% on ARC-AGI-1 benchmark

teortaxesTex · x · 2026-08-09

DeepSeek V4 Flash 0731 demonstrates strong performance on the ARC-AGI benchmark. At max effort, the model scores 89.0% on the ARC-AGI-1 Semi-Private set ($0.02 per task) and 61.4% on ARC-AGI-2 Semi-Private ($0.04 per task). The original poster noted the irony of AI evaluation timelines, remarking that by the time the highly intensive ARC-AGI-3 evaluation is completed, the model being tested will likely already be obsolete.

Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance, Costing a Quarter of GPT-5.6(9 posts)→

Original post →

More from Models

Models channel →