Frontier models fail 96% of simple 2D mazes in Long-Horizon test

thebigbigbuddha · reddit · 2026-09-02

MultiNet 2.0 tested frontier reasoning models on simple 2D mazes to study long-horizon interactive tasks. While trivial for humans, the mazes require continuous perception, reasoning, acting, and progress tracking. Out of 150 evaluation runs, the models solved only 6. The study aims to break down failure modes.

Original post →

More from AGI Musings

AGI Musings channel →