DeepSeek V4.1 Flash hits 72.9% on ARC-AGI-2 at $0.13/task, costing 250% more
teortaxesTex · x · 2026-10-07
ARC Prize published DeepSeek V4.1 Flash scores on ARC-AGI (Verified): 72.9% on ARC-AGI-2 ($0.13/task) and 94.5% on ARC-AGI-1 ($0.07/task).
That beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but at roughly 250% higher cost per task.
Commenter teortaxesTex notes it still fails to match the earlier, smaller Dots3-preview on either performance or cost, though he sees it as strong vindication of DeepSeek's RL scheme.
More from Models
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07
- Unverified rumor suggests Qwen4 Flash is a 400B parameter model, comparable to GLM 5.3 Flash — EAccelerate_42 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- Mistral Large 4.0 weights reportedly landing at end of October — cpldcpu · 2026-10-07
- Claude's "reasoning extraction" guardrail blocks users from seeing its thinking, and they're not happy — StewartalsopIII · 2026-10-07
- Grok's quirk: it says 'No.' then argues your point better than you did — gandamu_ml · 2026-10-07