GPT-6 'Astra' burns 60+ hours in 3D spatial reasoning eval, scores just 14%

patience_cave · x · 2026-09-07

Independent eval MazeBench tested "GPT-6 Astra" in a 3D open-world spatial reasoning benchmark. Astra spent over 60 hours in the environment and still finished with only 14% — the best result on the eval so far, but far from solving its hardest challenges.

Related event: GPT-6 Astra Sets MazeBench Record at 14% but Still Fails 3D Spatial Reasoning(5 posts)→

Original post →

More from Models

Models channel →