MazeBench launches as open long-horizon agent benchmark

MazeBench tests long-horizon planning and visuospatial reasoning in a 3D open world, with non-Python agents scoring just 1%; the fully open-source benchmark lets any model and harness compete.

2026-08-21 ~ 2026-08-21 · 2 related posts