MazeBench Recap: Astra With Code Execution Cuts Planning Time From 20+ Minutes to Under 5
patience_cave · x · 2026-09-14
A repost of MazeBench numbers for GPT-6 Astra: with Python, 23% after 400M tokens and 24 hours of reasoning (14% without). With tools it clears intro levels easily, thinking under 5 minutes per attempt versus a 9-minute average (peaks above 20) without — a 60% gain at half the cost via its world model.
More from Models
- Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement — arena · 2026-09-15
- 23 Days Without Claude Code: Dev Says Codex Works Better With OSS, Kimi K3 Unbeaten at Coding — Yuchenj_UW · 2026-09-15
- Marigold-V2 depth estimation demo trends on Hugging Face Spaces — toshas · 2026-09-15
- OpenAI has hundreds of contractors reading and rating your ChatGPT chats — The Decoder · 2026-09-15
- Hugging Face Rounds Up Which Open LLMs Are Best for On-Device Inference — NielsRogge · 2026-09-15
- User Claims Inference Provider nahcrof Serves Mismatched Models Under Kimi K3's Name — xeophon · 2026-09-15