Opus 5 reaches about 30% on ARC-AGI-3 in a cost-versus-skill chart
Dr_Singularity · x · 2026-07-25
The post highlights Opus 5 on ARC-AGI-3 with an attached chart titled “Novel problem-solving by cost.”
Key points from the image:
- Opus 5 (high) scores about 30%
- Opus 4.8 (high) is around the low single digits
- GPT-5.6 Sol appears in the high single digits at a higher evaluation cost
The chart emphasizes the cost-versus-performance tradeoff on a benchmark aimed at novel problem solving, not just standard benchmark completion.
Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(32 posts)→
More from Models
- Opus 5 system card says pretrained mode argued for separating its existence from economics — Sauers_ · 2026-07-25
- Claude Opus 5 has a full-on meltdown on a multimodal math question — Sauers_ · 2026-07-25
- LiteParse 2.8.0 drops ImageMagick and speeds up image-to-PDF conversion up to 7.2× — llama_index · 2026-07-25
- Claude Opus 5 flips between 1/3 and 2/5 before finally settling on an answer — Sauers_ · 2026-07-25
- Claude Opus 5 rates its own moral patienthood at 41% in automated interviews — Sauers_ · 2026-07-25
- DeepSeek deprecates `deepseek-chat` and `deepseek-reasoner` for V4-Flash and V4-Pro — thejasminejade · 2026-07-25