Claude Fable 5.1 hits 90% on ARC-AGI-2 semi-private at $4.49 per task
geoffwolfe · x · 2026-09-19
- ARC Prize's page shows Claude Fable 5.1 scoring 97.5% on ARC-AGI-1 Semi-Private at $1.40/task and 90.0% on ARC-AGI-2 Semi-Private at $4.49/task at max reasoning effort.
- Across five effort levels, ARC-AGI-2 scores scale from 78.3% (Low) to 90.0% (Max), showing consistent gains with inference budget; ARC-AGI-1 is nearly saturated.
- The verified leaderboard also lists variants from DeepSeek V4, Gemini 3.x, GPT-5.6/6, Grok 4.x and Kimi K3.
- Commenters note ARC-AGI-2 has become a score-and-cost benchmark, with remaining signal concentrated in the hard tail that still separates reasoning effort, consistency, and inference budget.
More from Models
- Dev suggests AI companies should refund API credits on agent refusals — GabGarrett · 2026-09-19
- Dev claims Codex users are 'getting scammed' over usage terms — AIFlow_ML · 2026-09-19
- Step 5 Preview Quietly Debuts on AA: Score 44, 1M Context, $1/1M Input — teortaxesTex · 2026-09-19
- Developers say Codex is 'not sustainable' under current usage limits — AIFlow_ML · 2026-09-19
- Cognition ships SWE-2 coding model: 1 point behind Fable 5.1 at 64% less cost — AxSaucedo · 2026-09-19
- Jev Hits 36M Views in 2 Days, Community Ships 6 Open Clones — Latent Space · 2026-09-19