GPT-6 Astra Reportedly Saturates ARC-AGI-3 with Fewer Steps Than Humans

GPT-6 Astra delivered a breakthrough result on the ARC-AGI-3 benchmark: according to ARC Prize's official blog, the model not only saturates the benchmark but completes tasks in fewer steps than the average human. If confirmed, this marks the fall of a benchmark that had stumped every frontier model, and reignites the debate over "scaling vs. symbolic reasoning."

Confirmed

Unconfirmed

Why it matters

2026-09-04 ~ 2026-09-04 · 5 related posts

Primary sources