GPT-6 Astra Scores Just 23% on MazeBench After 400M Tokens and 24 Hours of Reasoning
patience_cave · x · 2026-09-14
Running GPT-6 "Astra" with Python code execution on the MazeBench 3D environment benchmark consumed 400 million tokens and 24 hours of reasoning for a final score of just 23%; without Python, Astra drops to 14%.
The author notes Astra understands basic spatial concepts like walking beneath archways, but becomes unreliable in longer tunnels — 3D spatial reasoning still has plenty of room to grow.
More from Models
- Claude Max and Codex tiers are creating a computing power gap that locks out $20-budget newcomers — IndraVahan · 2026-09-15
- OpenAI CFO says 80% Luna price cut drove a 10X usage surge — Beth_Kindig · 2026-09-15
- Palantir, Nvidia, Booz Allen restrict Anthropic's Fable over data-retention concerns — dotey · 2026-09-15
- K2 Horizon 7B ranks between Qwen3.6 27B and 35BA3b on AA Intelligence Index — Uncle___Marty · 2026-09-15
- Users claim GPT-6 Astra got dumber and burns 5-hour quota in 5 minutes — Vibe_Mint · 2026-09-15
- Post-training compute shift makes distillation of frontier models nearly impossible to prevent — maksym_andr · 2026-09-15