Anthropic Opus 5.5 scores 93.3% on ARC-AGI-2 at 80% lower eval cost, ARC Prize confirms
burny_tech · x · 2026-09-23
- ARC Prize published verified Opus 5.5 results: 93.3% on ARC-AGI-2 ($0.41/task) and 98.5% on ARC-AGI-1 ($0.16/task).
- Compared to Opus 5, that's +2.9 points on v2 and +1.0 on v1, at roughly 80% lower evaluation cost.
- Official benchmark numbers from the ARC Prize account itself.
More from Models
- One RL Run Cost 130 Hours, 75B Tokens and $2.6M — the Real Price of Scaling RL — burny_tech · 2026-09-23
- Qwen 3.8 27B local coding session runs 3 days on one RTX 4090, then spews endless slashes — Tiny-Entertainer-346 · 2026-09-23
- Claude adds banked usage limit reset button on web and desktop — airesearch12 · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23