DeepSeek V4 Pro on ARC-AGI: Matches Flash Score but with Higher Params

teortaxesTex · x · 2026-08-31

DeepSeek V4 Pro scored 90.5% on ARC-AGI-1 ($0.18/task) and 61.3% on ARC-AGI-2 ($0.60/task). Comparisons show it is precisely 0.1% lower than Flash-0731, with similar performance across reasoning levels. Analysis questions the value of the 5.6x total and 3x active parameters, as all updated V4 models perform similarly on third-party evals.

Original post →

More from Models

Models channel →