Open benchmarking challenge says Opus 5’s ARC-AGI-3 leap doesn’t carry over

rbhar90 · x · 2026-07-25

The post argues for more open benchmarking, saying it no longer trusts strong reasoning claims from labs because too much money is now tied to the narrative.

It quotes a comparison around Opus 5 on ARC-AGI-3 versus the authors’ held-out Witness suite:

The broader point is that benchmark scores can overstate reasoning gains if the evaluation set and the real task distribution diverge.

Related event: Opus 5's Record ARC-AGI-3 Score Sparks Cheating and Overfitting Allegations(11 posts)→

Original post →

More from Models

Models channel →