GPT-6 Astra struggles on MazeBench 3D spatial reasoning benchmark

Blogger patiencecave tested GPT-6 Astra on MazeBench, a self-built 3D open-world spatial reasoning benchmark: the model ran for over 60 hours on the evaluation and scored only 14%. While low in absolute terms, this is the strongest performance on MazeBench to date—and achieved without code execution, beating GPT-5.6 Sol's 13% even with Python code execution enabled—leading the blogger to see it as a generational leap in capability.

Confirmed

Why it matters

Even the strongest new-generation models still have structural weaknesses in long-horizon, true 3D spatial reasoning, suggesting that "top-down-view" world understanding remains a bottleneck for current models; meanwhile, the sharp drop in token consumption from batch planning also demonstrates how agent benchmark design affects cost. The blogger notes there is still a large gap to passing MazeBench's harder challenges.

2026-09-07 ~ 2026-09-07 · 7 related posts

Primary sources

1 near-duplicate retellings: basedjensen