GPT-6 Astra Nearly Doubles GPT-5.6 Sol on MazeBench, Wins Even Without Code Execution
patience_cave · x · 2026-09-14
On the MazeBench 3D spatial reasoning benchmark, GPT-6 "Astra" with code execution nearly doubled the score of GPT-5.6 "Sol" with code execution. Surprisingly, Astra without any tools still beat GPT-5.6 Sol with code execution.
The author calls Astra a powerful agent, though earlier runs show it still struggles with longer tunnels in 3D environments.
Related event: MazeBench: GPT-6 Astra Scores 23% After 400M Tokens, 24 Hours(6 posts)→
More from Models
- OpenBMB ships MiniCPM5-2B, a local open model billed as best in its size class — thione · 2026-09-15
- Inception ships Mercury 2.5, a diffusion LLM claiming 40% higher intelligence at 1,107 tokens/sec — thione · 2026-09-15
- OpenAI launches GPT-Live-1 API with full-duplex voice, interruption handling and tool delegation — thione · 2026-09-15
- DeepSeek releases V4.1-Flash, a multimodal API model with faster inference and lower prices — thione · 2026-09-15
- Artificial Analysis Launches v1.1 Capability Indices; Claude Tops All Six Domains — ArtificialAnlys · 2026-09-15
- Forge launches with Arcee, Microsoft, Vercel to push open weight models to the frontier — inkko44 · 2026-09-15