GPT-6 Astra scores 14% on no-Python MazeBench, 7x Claude Fable 5.1's 2%
rohanpaul_ai · x · 2026-09-08
- MazeBench is a long-horizon spatial reasoning eval: 100 gems hidden across 200+ rooms, requiring view rotation, box manipulation, error recovery, and plans exceeding 100 moves.
- The no-Python track is much harder since code-enabled agents can reverse-engineer physics and run solvers instead of carrying the spatial plan themselves. GPT-6 Astra scored 14%, 7x Claude Fable 5.1's 2%.
- Astra reportedly spent 60+ hours in the 3D open-world eval, holding long plans in context; remaining failures on genuinely 3D puzzles suggest gains are mainly in sustained planning, not full spatial understanding.
- Note: benchmark and scores come from a personal tweet; no official confirmation yet.
More from Models
- X users accuse OpenAI of pressuring a renowned mathematician into academic fraud — S_Conradi · 2026-09-09
- Vercel Labs' gpu-lexer: a 27.5KB model does GPU-powered syntax highlighting in browser — shadcn · 2026-09-09
- Magic's roadmap: long-context RL, latent-knowledge alignment, then a model release — magicailabs · 2026-09-09
- Thomson Reuters' frontier-competitive legal model Thomson trained for just $450K with curated data — schwarzjn_ · 2026-09-09
- Scoop: Anthropic reportedly prepping fresh Fable-class pretrain to launch before September IPO — teortaxesTex · 2026-09-09
- User claims GPT-6 'Astra' is a step-function leap in generality, effectively AGI — brandon_galang · 2026-09-09