Open benchmarking challenge says Opus 5’s ARC-AGI-3 leap doesn’t carry over
rbhar90 · x · 2026-07-25
The post argues for more open benchmarking, saying it no longer trusts strong reasoning claims from labs because too much money is now tied to the narrative.
It quotes a comparison around Opus 5 on ARC-AGI-3 versus the authors’ held-out Witness suite:
- Opus 5 reportedly scores 30% on ARC-AGI-3, far ahead of prior models.
- On Witness, the jump does not transfer: Opus 5 lands at 43.4 ± 3.2, essentially tied with kimi-k3 (42.8 ± 1.9) and Fable-5 (43.8 ± 9.7).
- It is better than Opus 4.8 (34.8), but the authors say this is not a generational leap.
- The traces suggest Opus 5 already “knows” this class of puzzle: it states hidden rules before acting and then plays a byte-identical optimal solution in all 5 seeds at temperature 1.0, with zero exploration.
The broader point is that benchmark scores can overstate reasoning gains if the evaluation set and the real task distribution diverge.
More from Models
- Claude Opus 5 reportedly nails a snowboarder test in one shot — rohanpaul_ai · 2026-07-27
- GPT 5.6 Pro reportedly beats Work/Codex High and Extra High — latticecut · 2026-07-27
- A developer says Codex is now their coding environment inside ChatGPT — kevinkern · 2026-07-27
- Looped Transformers work best from scratch, with two passes emerging as the sweet spot — jm_alexia · 2026-07-27
- Nebius readies for tomorrow’s Kimi launch, with a bigger team and hiring stakes — demian_ai · 2026-07-27
- Claude Opus 5 ranks third on VoxelBench, just 30 Elo behind the leader — legit_api · 2026-07-27