MIT paper resurfaces: SOTA multimodal models failed dramatically on BABA puzzles

moschles · reddit · 2026-10-09

A Redditor resurfaces a 2024 ICML paper by MIT and Virginia Tech researchers, "BABA-is-AI": testing GPT-4o, Gemini-1.5-Pro and Gemini-1.5-Flash — then state-of-the-art multimodal LLMs — it found they fail dramatically when generalization requires manipulating and combining the rules of the game. The repo is nacloos/baba-is-ai on GitHub (arXiv:2407.13729).

The question raised: with today's tera-parameter agentic swarms already acing ARC-AGI-3 and FrontierMath tier 4, are these small key-door puzzles now trivially solvable? The author leans toward yes, given a suitable harness; but if not, the paper's importance has compounded, and it should be relayed to Francois Chollet and the ARC Foundation as a candidate benchmark for ARC-AGI-4.

Related event: Two-Year-Old BABA-is-AI Paper Still Stumps SOTA Models(2 posts)→

Original post →

More from Research

Research channel →