DeepSeek V4 Pro on ARC-AGI: Matches Flash Score but with Higher Params
teortaxesTex · x · 2026-08-31
DeepSeek V4 Pro scored 90.5% on ARC-AGI-1 ($0.18/task) and 61.3% on ARC-AGI-2 ($0.60/task). Comparisons show it is precisely 0.1% lower than Flash-0731, with similar performance across reasoning levels. Analysis questions the value of the 5.6x total and 3x active parameters, as all updated V4 models perform similarly on third-party evals.
More from Models
- Google closed Jeff Dean's account while Gemini 3.5 Pro remains 'soon' — ryanmerket · 2026-08-31
- User Complains GLM 5.3 Flash Overthinks Simple Prompts and Throws Errors — nahmanhuh · 2026-08-31
- Tested: Qwen3.8-Flash-Next Is Faster but Fakes Completion in Hard Tasks — trashacct383 · 2026-08-31
- Google AI Admits to Outputting Racist Content About Latinos — HelpfulQuestions · 2026-08-31
- Taalas demo shows 14,000 tokens/second generation speed — rohanpaul_ai · 2026-08-31
- Why do RLVR skills transfer? Lack of rigorous theory in model training — voooooogel · 2026-08-31