New Benchmark 'Baba is Harbor' Evaluates Multiple Models
pmigdal · reddit · 2026-07-17
The authors have open-sourced a benchmark named Baba is Harbor, which ports levels from the game *Baba Is You* into the Harbor agentic RL environment to evaluate model performance on game-like tasks. In Stage 0 and Stage 1 tests, Claude Fable 5 and GPT-5.6 Sol were able to solve most levels. Fable 5 was faster but still about 4 times slower than humans. The authors noted no obvious signs of memorization and shared full details on the pipeline, attempt counts, token usage, and total costs on their blog.
More from Research
- Draft paper uses Markov-chain eigenfunctions to build partitions and speed up sampling — michaelchchoi · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- GitHub repo adds lightweight ternary QAT for Prism-ML Bonsai models — terminoid_ · 2026-07-21
- Qdrant co-hosts a Munich meetup on search, retrieval, and agentic RAG on July 23 — qdrant_engine · 2026-07-21
- GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs — ai-sage · 2026-07-21
- Paper models Transformer components as stochastic geometry and tests five architectures — Zhihua Liang · 2026-07-21