MazeBench Recap: Astra With Code Execution Cuts Planning Time From 20+ Minutes to Under 5

patience_cave · x · 2026-09-14

A repost of MazeBench numbers for GPT-6 Astra: with Python, 23% after 400M tokens and 24 hours of reasoning (14% without). With tools it clears intro levels easily, thinking under 5 minutes per attempt versus a 9-minute average (peaks above 20) without — a 60% gain at half the cost via its world model.

Related event: MazeBench Tests GPT-6 Astra: Code Execution Dramatically Boosts 3D Spatial Reasoning(6 posts)→

Original post →

More from Models

Models channel →