GPT-6 without code execution beats GPT-5.6 with Python on MazeBench

patience_cave · x · 2026-09-07

Wrap-up of the MazeBench thread: GPT-6 Astra scored 14% without code execution, beating GPT-5.6 Sol's 13% even when the latter had Python available — a clear generational leap, yet still far from the eval's hardest challenges. The MazeBench leaderboard is public at mazebench.com.

Related event: GPT-6 Astra Sets MazeBench Record at 14% but Still Fails 3D Spatial Reasoning(5 posts)→

Original post →

More from Models

Models channel →