MazeBench: 400M tokens and 24 hours of reasoning scores just 23% for GPT-6 Astra

mhmazur · x · 2026-09-14

The MazeBench author tested GPT-6 Astra in a 3D maze environment: after 400 million tokens and 24 hours of reasoning, it scored just 23% — and only 14% without Python tooling, showing code tools matter greatly for spatial reasoning. Others praised the benchmark and analysis.

Related event: MazeBench Tests GPT-6 Astra: Code Execution Dramatically Boosts 3D Spatial Reasoning(6 posts)→

Original post →

More from Models

Models channel →