Anthropic’s Opus 5 hits 30% on ARC-AGI-3, but loses its edge on a held-out puzzle suite
NathanpmYoung · x · 2026-07-27
Anthropic’s Opus 5 is said to reach 30% on ARC-AGI-3, roughly 4× the previous best and 20× its predecessor Opus 4.8.
But a separate test on Witness, an in-house held-out suite of ARC-AGI-3-style interactive puzzle games, shows the gain does not transfer:
- Opus 5: 43.4 ± 3.2
- kimi-k3: 42.8 ± 1.9
- Fable-5: 43.8 ± 9.7
- Opus 4.8: 34.8
The critique says Opus 5 appears to already know this puzzle genre: on one classic game it states the hidden rules before making its first move and then produces the same optimal solution across all five seeds at temperature 1.0, suggesting zero exploration rather than a true generational leap.
More from Models
- Reddit argues Qwen should keep shipping capable 27B–397B models instead of 2T+ giants — Responsible_Fig_1271 · 2026-07-27
- A thread argues serious Opus 3 use becomes a superpower as models improve — repligate · 2026-07-27
- Model chart pits Opus 5, Sonnet 5 and GPT-5.6 against cost and coding scores — haider1 · 2026-07-27
- Sam Altman says the next six months of model progress will outpace the last two years — ycombinator · 2026-07-27
- Would a $5/month coding-model plan beat Claude and Codex at $20/month? — athsrva · 2026-07-27
- Poster says Kimi’s near-term outlook depends on a K3 base model release — _xjdr · 2026-07-27