GPT-6 Astra scores just 14% on MazeBench 3D spatial reasoning eval after 60+ hours

basedjensen · x · 2026-09-07

MazeBench, a 3D open-world spatial reasoning benchmark, put GPT-6 Astra through 60+ hours of testing with a final score of only 14%. Reposter Andrew Curran says he trusts this benchmark more than 99% of existing ones, noting capability progress is jagged — users increasingly see only the parts of the elephant relevant to their tasks.

Related event: GPT-6 Astra scores just 14% after 60-hour 3D spatial reasoning benchmark run(7 posts)→

Original post →

More from Models

Models channel →