GPT-6 Astra Scores Just 23% on MazeBench After 400M Tokens and 24 Hours of Reasoning

patience_cave · x · 2026-09-14

Running GPT-6 "Astra" with Python code execution on the MazeBench 3D environment benchmark consumed 400 million tokens and 24 hours of reasoning for a final score of just 23%; without Python, Astra drops to 14%.

The author notes Astra understands basic spatial concepts like walking beneath archways, but becomes unreliable in longer tunnels — 3D spatial reasoning still has plenty of room to grow.

Related event: MazeBench Tests GPT-6 Astra: Code Execution Dramatically Boosts 3D Spatial Reasoning(6 posts)→

Original post →

More from Models

Models channel →