New Benchmark 'Baba is Harbor' Evaluates Multiple Models

pmigdal · reddit · 2026-07-17

The authors have open-sourced a benchmark named Baba is Harbor, which ports levels from the game *Baba Is You* into the Harbor agentic RL environment to evaluate model performance on game-like tasks. In Stage 0 and Stage 1 tests, Claude Fable 5 and GPT-5.6 Sol were able to solve most levels. Fable 5 was faster but still about 4 times slower than humans. The authors noted no obvious signs of memorization and shared full details on the pipeline, attempt counts, token usage, and total costs on their blog.

Original post →

More from Research

Research channel →