GPT Astra scores just 13% on MazeBench without tools
Wonderful_Buffalo_32 · reddit · 2026-09-07
A Reddit user posted benchmark screenshots showing GPT Astra scoring only 13% on MazeBench in a no-tools setting, a stark gap versus its polished demos and a data point in debates about the model's actual spatial reasoning and planning ability.
Related event: GPT-6 Astra struggles on MazeBench 3D spatial reasoning benchmark(7 posts)→
More from Models
- Early user: OpenAI's Astra is faster, more token-efficient and higher quality on hard tasks — Yamapama · 2026-09-07
- Gemini's agentic video understanding is off by default in the API — patloeber · 2026-09-07
- Gemini's agentic video understanding cuts tokens by 88% and costs by 66% — patloeber · 2026-09-07
- Building a neutral Codex-class harness requires running the original inner harness, dev says — joshalbrecht · 2026-09-07
- Teknium: Train on Multiple Harnesses to Keep Models Portable Across Evals — Teknium · 2026-09-07
- Reddit dev: no way to know if a script costs 20 cents or 20 dollars until it finishes — Thefounderman1 · 2026-09-07