MazeBench: New 3D Spatial Reasoning Benchmark Where Prior SOTA Agents Score Just 1%
patience_cave · x · 2026-09-04
A new benchmark called MazeBench targets 3D spatial reasoning, and it's brutal: previous SOTA agents score only 1%. The author invites models like Astra and Fable to attempt the maze, positioning it as a hard test of spatial navigation and long-horizon reasoning.
More from Research
- Nature review: AI now designs physics experiments that rival human inventions — _sathvikr · 2026-09-04
- Google maps the complete male fruit fly brain in connectomics milestone — ThePlanckDiver · 2026-09-04
- Symbolic world models strike back: mu0 beats pi0.5 with 1/100 of the data — furongh · 2026-09-04
- Sparse Readout Prism explains Logit-Lens scores via sparse features instead of tokens — Matteo He · 2026-09-04
- ARC-AGI's Kamradt: clever harnesses measure human intelligence, not models — GregKamradt · 2026-09-04
- Harvard RCT: students learn more in less time with an AI tutor than active-learning classes — ___Patrice___ · 2026-09-04