MazeBench goes free online: agents prefer top-down views and stall in 3D puzzles
patience_cave · x · 2026-10-02
MazeBench, a benchmark for visual spatial reasoning in a 3D open world, is now free to play online. The author's observation: agents master 2D puzzles but struggle in 3D — they prefer top-down views and only rotate the camera as a last resort, which quickly stifles progress.
Related event: MazeBench Shows AI Agents Struggle in 3D Puzzles(2 posts)→
More from Research
- Sasha Rush heads to COLM 2025, inviting chats on TTT, proofs and biased RL — srush_nlp · 2026-10-02
- CISPA Offers Imprecise Probabilistic ML Course Again, Free and Online — krikamol · 2026-10-02
- SoftServe's updates build on quadratic matrix equations, extending the authors' Batch-and-Match BBVI work — dianarycai · 2026-10-02
- SoftServe preprint brings scalable quasi-Newton optimization to deep learning, beating Adam, Muon and SOAP on ill-conditioned tasks — dianarycai · 2026-10-02
- DyRAD: radar novel view synthesis renders full range-azimuth-Doppler tensors for dynamic driving scenes — orlitany · 2026-10-02
- EMBL-EBI's saezlab open-sources Karenina, a framework for multi-dimensional biomedical AI evaluation — anshulkundaje · 2026-10-02