MazeBench launches as open long-horizon agent benchmark
MazeBench tests long-horizon planning and visuospatial reasoning in a 3D open world, with non-Python agents scoring just 1%; the fully open-source benchmark lets any model and harness compete.
2026-08-21 ~ 2026-08-21 · 2 related posts
- MazeBench benchmark tests long-horizon agents: 0% score without Python tools — JFPuget · 2026-08-21
- MazeBench is fully open source: any model with any harness can enter — xeophon · 2026-08-21