Maze Bench: GPT-6 Astra without code execution reportedly beats GPT-5.6 Sol with code
patience_cave · x · 2026-09-14
Maze Bench, a visual spatial reasoning benchmark in a 3D open world (playable free online), reportedly shows GPT-6 Astra with code execution nearly doubling GPT-5.6 Sol's score. Bigger surprise: GPT-6 Astra without code execution still beats GPT-5.6 Sol with code, suggesting strong agentic spatial reasoning. Note: GPT-6 is unannounced, treat as unverified.
Related event: MazeBench: GPT-6 Astra Scores 23% After 400M Tokens, 24 Hours(6 posts)→
More from Models
- Seattle Times and Newsday sue OpenAI and Microsoft over AI training on journalism — thione · 2026-09-15
- OpenBMB ships MiniCPM5-2B, a local open model billed as best in its size class — thione · 2026-09-15
- Inception ships Mercury 2.5, a diffusion LLM claiming 40% higher intelligence at 1,107 tokens/sec — thione · 2026-09-15
- OpenAI launches GPT-Live-1 API with full-duplex voice, interruption handling and tool delegation — thione · 2026-09-15
- DeepSeek releases V4.1-Flash, a multimodal API model with faster inference and lower prices — thione · 2026-09-15
- Artificial Analysis Launches v1.1 Capability Indices; Claude Tops All Six Domains — ArtificialAnlys · 2026-09-15