MazeBench cracked by a classic BFS solver; author admits difficulty is perception-based
patience_cave · x · 2026-09-03
The MazeBench author discusses findings with independent developer FakePsyho, who built an offline solver after getting source access: a separate room solver (an optimized BFS) plus glue logic forming a world solver, taking only a few hours of code optimization to score well — and finding no levels designed adversarially against solvers.
The author concedes MazeBench's difficulty is currently perception-based rather than solver-based, plans updates to the latter half of levels, and will invest more in level design before astra goes deeper. He also notes one user-designed level took sol xhigh + codex hours to solve even with source access, plus a blind maze-tunnel level; the human leaderboard's gem count is capped around 80 because final levels are still being built and quality-checked.
Related event: Developer Cracks MazeBench with a Simple BFS Solver(3 posts)→
More from Research
- Revera: Lean-verified POSIX regex engine with identical output in 6 languages — jedisct1 · 2026-09-03
- ChatGPT helps resolve 30-year-old stable forking conjecture in logic — Dr_Singularity · 2026-09-03
- Yale trio's new paper: mechanism design for AI agents with unknown alignment — Afinetheorem · 2026-09-03
- PufferLib 5.0 Self-Play Trains 10-Ship Duel in 5 Minutes on a $700 PC — yacineMTB · 2026-09-03
- Until Labs scales cryoprotectant discovery to 250,000+ candidate molecules with AI — lukaszkaiser · 2026-09-03
- Stanford's Michael Bernstein builds a "What-If Machine" for simulating decisions with AI — msbernst · 2026-09-03