MazeBench: 400M tokens and 24 hours of reasoning scores just 23% for GPT-6 Astra
mhmazur · x · 2026-09-14
The MazeBench author tested GPT-6 Astra in a 3D maze environment: after 400 million tokens and 24 hours of reasoning, it scored just 23% — and only 14% without Python tooling, showing code tools matter greatly for spatial reasoning. Others praised the benchmark and analysis.
More from Models
- Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement — arena · 2026-09-15
- 23 Days Without Claude Code: Dev Says Codex Works Better With OSS, Kimi K3 Unbeaten at Coding — Yuchenj_UW · 2026-09-15
- Marigold-V2 depth estimation demo trends on Hugging Face Spaces — toshas · 2026-09-15
- OpenAI has hundreds of contractors reading and rating your ChatGPT chats — The Decoder · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- User Claims Inference Provider nahcrof Serves Mismatched Models Under Kimi K3's Name — xeophon · 2026-09-15