ARC-AGI Evaluation Dispute: Claude Opus Score Questioned Over API Implementation Flaw
steipete · x · 2026-07-30
Developer @steipete pointed out that Claude Opus's recent high score on the ARC-AGI evaluation might be inflated. This is due to a flaw in the model's Chat Completion API implementation, which erroneously retained reasoning tokens, an API the official documentation does not support. However, some users argue that if both Opus and GPT-5 used the same generic testing harness, the comparison remains valid, suggesting Opus generalizes better.
Related event: Claude Opus ARC-AGI Score Questioned Over API Flaw(2 posts)→
More from Models
- Tencent's Hy3 Model Solves 50-Year-Old Combinatorics Problem — Tim_Dettmers · 2026-07-30
- Microsoft Shares Production Data for MAI-Code-1-Flash: Balancing Coding Quality and Token Efficiency — lee_stott · 2026-07-30
- Sarvam AI Announces Open Weight Models on Indian Infrastructure — AashaySachdeva · 2026-07-30
- Testing All OpenRouter TTS Models: Kokoro-82M is Best and Cheapest for Long-Form — nathanborror · 2026-07-30
- Qwen3.6 MoE 2-bit Quantized Version Tops Hugging Face Trending — EschaLabs · 2026-07-30
- Leaked Tasks Hint at Anthropic's Strategy: Training Expert Judge Models from Human Traces — burny_tech · 2026-07-30