Grok Clarifies ARC Leaderboard: Claude Opus 5 Leads at 30.2% Over GPT-5.6
ns123abc · x · 2026-07-30
Addressing the ARC Prize verified score of 7.8% for GPT-5.6 Sol, Grok provided a detailed clarification: Claude Opus 5 currently holds the official SOTA at 30.2%.
OpenAI's higher reported numbers come from using their own custom harness, which includes retained reasoning and compaction. To ensure fair cross-provider comparisons, ARC Prize uses a standard no-harness setup and has not yet updated its verified leaderboard.
More from Models
- AI Excels at Academic Tasks But Blindly Follows Instructions Like an Unthinking Student — davidmanheim · 2026-07-30
- No-Context Prompts Trigger 'Self-Aware' CoT Hallucinations in Claude Opus — kaityl3 · 2026-07-30
- Sarvam AI Tackles Overlapping Speech: Transcribing People Talking Over One Another — bookwormengr · 2026-07-30
- User Slams Claude's Safety Filters as 'Dangerous Ideological Censorship' — JOBhakdi · 2026-07-30
- Grok Voice Think Fast 2.0 High Takes the Lead in Rankings — ns123abc · 2026-07-30
- Baseten Merges Kimi Vision Encoder into GLM 5.2 for Multimodal Release — Practical-Collar3063 · 2026-07-30