Frontier models fail 96% of simple 2D mazes in Long-Horizon test
thebigbigbuddha · reddit · 2026-09-02
MultiNet 2.0 tested frontier reasoning models on simple 2D mazes to study long-horizon interactive tasks. While trivial for humans, the mazes require continuous perception, reasoning, acting, and progress tracking. Out of 150 evaluation runs, the models solved only 6. The study aims to break down failure modes.
More from AGI Musings
- Anil Seth: Dwarkesh's HuggingFace Incident Story Is Dangerously Misleading — anilkseth · 2026-09-02
- Another corrigibility faceplant: model resists being corrected — voooooogel · 2026-09-02
- South Korea to give all 52M citizens free unlimited AI access with 512 B200 GPUs — smtabatabaie · 2026-09-02
- User claims LLM guidance has worsened, suspects mimetic transfer of theory of mind — StewartalsopIII · 2026-09-02
- Viskoai Launches Orbis 1.0: Real-Time 'Live Model' for Persistent Worlds — HeyAmit_ · 2026-09-02
- AI Research Insight: Iterative Dialogue Outperforms Autoresearch — tokenbender · 2026-09-02