Opus 5 reaches about 30% on ARC-AGI-3 in a cost-versus-skill chart
Dr_Singularity · x · 2026-07-25
The post highlights Opus 5 on ARC-AGI-3 with an attached chart titled “Novel problem-solving by cost.”
Key points from the image:
- Opus 5 (high) scores about 30%
- Opus 4.8 (high) is around the low single digits
- GPT-5.6 Sol appears in the high single digits at a higher evaluation cost
The chart emphasizes the cost-versus-performance tradeoff on a benchmark aimed at novel problem solving, not just standard benchmark completion.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11