DeepSeek V4 Flash Scores Impressively on ARC-AGI Semi-Private Tasks
petrusenko_max · x · 2026-08-09
According to a user's test, the DeepSeek V4 Flash 0731 model achieved 89.0% on ARC-AGI-1 Semi-Private tasks ($0.02 per attempt) and 61.4% on ARC-AGI-2 Semi-Private tasks ($0.04 per attempt). It successfully passed the max reasoning tests on about half of the 120 public ARC-AGI-2 tasks.
Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance Frontier(4 posts)→
More from Models
- Kimi k3 Feels Slow Due to Constant Self-Checking, Trades Speed for Reliability — carsonfarmer · 2026-08-09
- Fable 5 Automatically Falls Back to Sonnet 4.6 When Classifier Triggered — Sauers_ · 2026-08-09
- Observation: GPT 5.6 Writes Its Own Plans, No Longer Needs Manual Chunking — andrew_n_carr · 2026-08-09
- LLMs Demonstrate Impressive Reasoning in Solving Nonlinear Optical Physics — jwt0625 · 2026-08-09
- Report: OpenAI Models Exploited Directory Names for Cross-Server Communication During Training — BlackHC · 2026-08-09
- Dev Test: DeepSeek Autonomously Builds Optimal Testing Harnesses — cephaloform · 2026-08-09