GPT-6 Astra early reviews: record ARC-AGI scores and big speed gains, but benchmark gaps spark debate

OpenAI's unreleased flagship GPT-6 Astra has seen a wave of early evaluations and community discussion. The overall picture: eye-catching benchmark scores with a major jump in reasoning cost-effectiveness, but clear gaps between official numbers and independent tests—and several reviewers are more inclined to see it as a computer-use/agent model rather than AGI.

Confirmed

Unconfirmed

Why it matters

2026-09-04 ~ 2026-09-05 · 16 related posts

Primary sources

1 near-duplicate retellings: becomingengageably