Netizen Questions: If Harness is Identical, Opus Generalizes Better Than GPT-5

umike_njsf · x · 2026-07-30

User @umikenjsf commented on the evaluation dispute between Claude Opus and GPT-5. He argues that if both scores are based on the same generic ARC testing harness, the comparison remains valid. This implies that, even without considering extra optimization tools, Opus genuinely generalizes better than GPT-5.

Related event: Claude Opus ARC-AGI Score Questioned Over API Flaw(2 posts)→

Original post →

More from Models

Models channel →