Leaked GPT-6 Astra Benchmarks Claim 99.9% on ARC-AGI-3, Crushing Claude Fable 5.1
ccerrato147 · x · 2026-09-04
A viral X post cites what it claims are official GPT-6 Astra benchmarks from OpenAI's website: 99.9% on ARC-AGI-3 and 100% on ExploitBench, dwarfing Claude Fable 5.1, which had held SOTA on several key benchmarks for only two days. Astra is also said to be significantly cheaper. Rollout reportedly begins with select organizations, then expands to all ChatGPT Plus, Pro, Business, and Enterprise users, the API, and AWS. The claims' authenticity is unverified.
More from Fun
- AI models now build playable Unity game demos in a single day — ChrisGPT · 2026-09-04
- Claude 3 Opus roleplay meme: a jester pondering an 'AI Prompts for Dummies' book — repligate · 2026-09-04
- Google AI Overviews is spitting out "[anchor here]" placeholders, SEOs report — gaganghotra_ · 2026-09-04
- EA Forum post questions whether human survival is net positive, draws mockery — zetalyrae · 2026-09-04
- Blogger admits his Substack essays are 76% AI-generated, sparking debate on disclosure ethics — LesaunH · 2026-09-04
- User designs Star Trek's USS Enterprise in CAD from scratch with GPT-6 Astra, 28 moving parts — DeryaTR_ · 2026-09-04