Opus 5 reaches 30.2% on ARC-AGI 3 as critics question the benchmark
ChrSzegedy · x · 2026-07-25
Opus 5 scores 30.2% on ARC-AGI 3, as ARC-AGI gets dismissed as an AGI proxy
Chris Szegedy says ARC-AGI has nothing to do with AGI, in response to a post claiming “AGI is near.”
- The thread cites Opus 5 scoring 30.2% on ARC-AGI 3.
- The argument is less about the score itself and more about whether ARC-style benchmarks meaningfully measure AGI.
- It is a concise data point plus a broader benchmark-validity critique.
Related event: Claude Opus 5's High ARC-AGI-3 Score Sparks Cheating Allegations(7 posts)→
More from AGI Musings
- Jensen Huang says heavier AI use can still mean more hiring — CodeByPoonam · 2026-07-25
- X debate says AI reviews could outclass many NeurIPS reviewers by 10x to 100x — peter_richtarik · 2026-07-25
- The next AI battle may be about memory lock-in, not benchmarks — VraserX · 2026-07-25
- Reuters asks whether chatbots will turbocharge cyberattacks and reshape security — wschroll · 2026-07-25
- A 149-page survey says long-horizon agents depend on harnesses, not just bigger models — 机器之心 · 2026-07-25
- A short AI-community take says better models ship faster when guardrails take a back seat — victor_explore · 2026-07-25