Maze Bench: GPT-6 Astra without code execution reportedly beats GPT-5.6 Sol with code

patience_cave · x · 2026-09-14

Maze Bench, a visual spatial reasoning benchmark in a 3D open world (playable free online), reportedly shows GPT-6 Astra with code execution nearly doubling GPT-5.6 Sol's score. Bigger surprise: GPT-6 Astra without code execution still beats GPT-5.6 Sol with code, suggesting strong agentic spatial reasoning. Note: GPT-6 is unannounced, treat as unverified.

Related event: MazeBench: GPT-6 Astra Scores 23% After 400M Tokens, 24 Hours(6 posts)→

Original post →

More from Models

Models channel →