GPT-6 Astra scores just 14% on MazeBench 3D spatial reasoning eval after 60+ hours
basedjensen · x · 2026-09-07
MazeBench, a 3D open-world spatial reasoning benchmark, put GPT-6 Astra through 60+ hours of testing with a final score of only 14%. Reposter Andrew Curran says he trusts this benchmark more than 99% of existing ones, noting capability progress is jagged — users increasingly see only the parts of the elephant relevant to their tasks.
More from Models
- Chess benchmarks questioned: a model could just download Stockfish and crush Magnus — iruletheworldmo · 2026-09-08
- Claude Fable 5.1 system prompt reveals why Claude 5 models wrote so badly — dbreunig · 2026-09-08
- Frontier models now lead local models by only about 6 months — Anxious_Current2593 · 2026-09-08
- Leaked ChatGPT debug stream reveals one prompt triggers 18 hidden queries and 50 engine calls — metehan777 · 2026-09-08
- Anthropic user wants normal-talking Opus back despite praising Fable 5.1 — james_mtc · 2026-09-07
- User notes Astra follows strict skill rules less reliably than Sol — petergyang · 2026-09-07