GPT-6-Astra benchmark leaks: 98.6% on ARC-AGI-3, perfect score on ExploitBench
teropa · x · 2026-09-04
- X user scaling01 posted benchmark results for GPT-6-Astra (no official blog yet):
- ARC-AGI-3: 98.6%
- FrontierMath Tier 4 v2: 97.6%
- DeepSWE v1.1: 74.1%
- ExploitBench: 100%
If accurate, the model nearly saturates abstract reasoning and frontier math benchmarks; numbers remain unverified pending official confirmation.
Related event: GPT-6 Astra Launches with Benchmark Leaks: ARC-AGI-3 Hits 98.6%(41 posts)→
More from Models
- Epoch AI: GPT-6 Astra sets ECI record of 169, tops math and continual learning benchmarks — NathanpmYoung · 2026-09-04
- GPT-6 Astra system card: no-CoT time horizon up ~10x over GPT-5.6 Sol, UK AISI finds — scaling01 · 2026-09-04
- Every's Vibe Check: GPT-6 Astra Is a Big Upgrade, but Anthropic's Fable Still Has Better Product Instincts — every · 2026-09-04
- GPT-6 Astra Nukes ARC-AGI-3: Score Jumps from 8% to 63%, 98.6% with Adapter — haider1 · 2026-09-04
- Sam Altman Officially Launches GPT-6 Astra, Claiming Best-in-World Computer Use and Coding — eyishazyer · 2026-09-04
- Early user: GPT-6 Astra rebuilt Apple Park in Blender from just images — BLUECOW009 · 2026-09-04