MazeBench is a 3D benchmark that leaves today’s agents stuck in the opening rooms
patience_cave · x · 2026-07-28
- MazeBench is introduced as a 3D open-world benchmark for long-term planning and visual-spatial reasoning.
- It contains hundreds of rooms and puzzles; the authors say current agents cannot get past the initial levels.
- A follow-up note adds that agents must traverse more than 200 rooms to find 100 hidden gems via Sokoban-like box-pushing puzzles, with difficulty ramping quickly.
- The team also reports that models were tested in native harnesses with camera/player tools, yet performance stayed extremely poor across image, ASCII, and JSON representations, topping out at 1%.
- To prevent endless loops, they used a novelty score to stop agents when they began circling.
Related event: MazeBench: New 3D Benchmark Leaves AI Agents Stuck at Level One(2 posts)→
More from coding & agent
- Eve adds shadcn registry support for browser agent extensions — shadcn · 2026-07-28
- Agent Mini packs a usable local AI agent into about 3,000 lines of Python — Lordrovks · 2026-07-28
- A simple tool-call fingerprint can stop many LLM agent loops, but not all of them — Future_AGI · 2026-07-28
- AI Tool Robin Reduces Dark Web Research to 30 Minutes — tom_doerr · 2026-07-28
- The real bottleneck in an Excel + LLM system is preserving business knowledge — Tired40s · 2026-07-28
- Four models tried the same game-building prompt, and Opus 5 looked finished — victor_explore · 2026-07-28