BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest

moschles · reddit · 2026-10-09

A 2024 ICML paper (BABA-is-AI, MIT + Virginia Tech) showed GPT-4o and Gemini-1.5 models "fail dramatically" when generalization requires manipulating and combining game rules. The poster asks whether tera-parameter agentic swarms that now ace ARC-AGI-3 and FrontierMath tier 4 can solve these small key-door puzzles — and if not, the paper's importance has only compounded. They suggest relaying it to Francois Chollet and the ARC Foundation as a candidate benchmark for ARC-AGI-4. Code and paper (arXiv 2407.13729) are public.

Related event: Two-Year-Old BABA-is-AI Paper Still Stumps SOTA Models(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →