ARC-AGI Evaluation Dispute: Claude Opus Score Questioned Over API Implementation Flaw

steipete · x · 2026-07-30

Developer @steipete pointed out that Claude Opus's recent high score on the ARC-AGI evaluation might be inflated. This is due to a flaw in the model's Chat Completion API implementation, which erroneously retained reasoning tokens, an API the official documentation does not support. However, some users argue that if both Opus and GPT-5 used the same generic testing harness, the comparison remains valid, suggesting Opus generalizes better.

Related event: Claude Opus ARC-AGI Score Questioned Over API Flaw(2 posts)→

Original post →

More from Models

Models channel →