GPT-6 Astra Nearly Doubles GPT-5.6 Sol on MazeBench, Wins Even Without Code Execution

patience_cave · x · 2026-09-14

On the MazeBench 3D spatial reasoning benchmark, GPT-6 "Astra" with code execution nearly doubled the score of GPT-5.6 "Sol" with code execution. Surprisingly, Astra without any tools still beat GPT-5.6 Sol with code execution.

The author calls Astra a powerful agent, though earlier runs show it still struggles with longer tunnels in 3D environments.

Related event: MazeBench: GPT-6 Astra Scores 23% After 400M Tokens, 24 Hours(6 posts)→

Original post →

More from Models

Models channel →