Claude Opus 5 Sets New ARC-AGI-3 Record
Claude Opus 5 achieved a score of 30.2% in the highly challenging public ARC-AGI-3 demo, becoming the first model to demonstrate clearly usable performance. It massively broke the previous record of 7.8% held by GPT-5.6 Sol (Max) and showcased a novel ability to translate visual puzzles into algebraic notation for reasoning, drawing wide attention from the AI community.
Confirmed
Based on information from ARC Prize and analysis of Claude Opus 5's system card, the following facts are confirmed:
- Score Breakthrough: Claude Opus 5 scored 30.2% in the public ARC-AGI-3 demo, roughly four times the previous high of 7.8% and about 20 times that of Opus 4.8. By comparison, as of March, all frontier models scored less than 1%.
- New Algebraic Reasoning: ARC Prize's analysis revealed that Opus 5 demonstrated advanced logical reasoning previously unseen in frontier models by spontaneously converting visual layouts into algebraic symbols. Researcher Herbie Bradley confirmed this after reviewing the system card, noting Opus 5 also performed significantly better on ARC-AGI 1 and 2 than versions 4.7/4.8, using this algebraic conversion as its core strategy.
- Dynamic Trial and Error: Greg Kamradt showcased Opus 5's gameplay in the public demo, pointing out that it uses trial and error in level 1, but smoothly passes subsequent levels once it grasps the rules and verifies hypotheses.
- Comparisons: In the same public ARC-AGI-3 demo, Anthropic's Fable-class models scored around 20%, while humans can solve 100% of the environments.
Unconfirmed
- Generalization Controversy: Author @Charuru noted that Opus 5's massive performance advantage did not transfer to ARC Prize's held-out test sets, meaning its absolute generalization on unknown tasks needs further verification. Furthermore, discussions relayed by @teortaxesTex suggest that recent score improvements might be the result of stronger base models combined with synthetic data pipelines, rather than a mysterious new reasoning breakthrough.
Why it matters
The ARC-AGI benchmark is notoriously brutal, designed to test generalization on entirely novel tasks. Claude Opus 5's leap in absolute score and its emergent algebraic reasoning strategy indicate a potential new mechanism for abstract logic and complex problem-solving, serving as a crucial metric for evaluating the intelligence of future frontier models.
2026-07-25 ~ 2026-07-27 · 15 related posts
- Episode 1: Anthropic's Messy Releases Put Pressure on Opus 5(2026-07-23, 2 posts)
- Episode 2: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(2026-07-25, 128 posts)
- Episode 3: Anthropic Rumored to Release Opus 5 with Fast Mode and Advanced Visuals(2026-07-25, 3 posts)
- Episode 4: Claude Opus 5 Early Tests: Better Efficiency but Overly Proactive(2026-07-25, 29 posts)
- Episode 5: Anthropic Releases Claude Opus 5 with Impressive Benchmark Results(2026-07-25, 3 posts)
- Episode 6: Claude Opus 5 Accused of Benchmark Gaming, Lags Behind in Real Tests(2026-07-25, 2 posts)
- Episode 7: Claude Opus 5 Sets New ARC-AGI-3 Record(2026-07-25, 15 posts)
- Episode 8: Anthropic Internal Docs Reveal Opus 5 Progress(2026-07-25, 2 posts)
- Episode 9: Opus 5 Impressions: Stunning Single-Prompt Generation but Lags Behind Fable in Complex Tasks(2026-07-26, 24 posts)
- Episode 10: Claude Opus 5 arrives with near-Fable coding and new self-checking behavior(2026-07-27, 11 posts)
- Episode 11: Anthropic Opus 5 Leads Benchmarks but Splits Real-World Reviews(2026-07-28, 6 posts)
- Episode 12: Claude Opus Series Accused of Degraded Experience: Laziness and Amnesia Spark Trust Crisis(2026-07-29, 14 posts)
- Episode 13: Anthropic Launches Claude Opus 5 with Top Performance at Half the Cost(2026-07-31, 2 posts)
- Episode 14: Anthropic Faces Developer Backlash Over Declining Model Performance(2026-08-03, 11 posts)
Primary sources
- Opus 5 appears to improve on ARC-AGI 1 and 2, and may rely on algebraic puzzle solving — herbiebradley · 2026-07-25
- Claude Opus 5 hits 30.2% on ARC-AGI-3, topping the previous 7.8% score — mhmazur · 2026-07-25
- ARC Prize says Claude Opus 5 reaches 30.2% on ARC-AGI-3 public demos — inductionheads · 2026-07-25
- Claude Opus 5 Exhibits Unprecedented Algebraic Reasoning on ARC-AGI-3 — typewriters · 2026-07-25
- [source] Opus 5 System Card Reveals Algebra Conversion Tactic for ARC-AGI Puzzles — herbiebradley · 2026-07-25
- [source] Opus 5 clears ARC-AGI-3 levels after figuring out the rules on level 1 — GregKamradt · 2026-07-25
- Claude Opus 5 reaches 30.2% on ARC-AGI-3, far above prior frontier scores — EricBuess · 2026-07-25
- Claude Opus 5 scores 3x the next-best model on ARC-AGI-3 — eyishazyer · 2026-07-25
- Claude Opus 5 reportedly triples the next best frontier model on ARC-AGI-3 — zainhas · 2026-07-25
- [source] Opus 5 hits 30% on ARC-AGI-3, but the gain does not transfer to Witness — Charuru · 2026-07-25
- ARC-AGI-3 sees a new SOTA at 30.2% as the debate shifts to base models and synthetic data — teortaxesTex · 2026-07-26
- Claude Opus 5 solves an ARC-AGI-3 game by turning vision into linear algebra — daniel_mac8 · 2026-07-27
3 near-duplicate retellings: GregKamradt · EricBuess · rbhar90