MIT paper resurfaces: SOTA multimodal models failed dramatically on BABA puzzles
moschles · reddit · 2026-10-09
A Redditor resurfaces a 2024 ICML paper by MIT and Virginia Tech researchers, "BABA-is-AI": testing GPT-4o, Gemini-1.5-Pro and Gemini-1.5-Flash — then state-of-the-art multimodal LLMs — it found they fail dramatically when generalization requires manipulating and combining the rules of the game. The repo is nacloos/baba-is-ai on GitHub (arXiv:2407.13729).
The question raised: with today's tera-parameter agentic swarms already acing ARC-AGI-3 and FrontierMath tier 4, are these small key-door puzzles now trivially solvable? The author leans toward yes, given a suitable harness; but if not, the paper's importance has compounded, and it should be relayed to Francois Chollet and the ARC Foundation as a candidate benchmark for ARC-AGI-4.
Related event: Two-Year-Old BABA-is-AI Paper Still Stumps SOTA Models(2 posts)→
More from Research
- Lancet study: patient-facing conversational AI holds up in real urgent care settings — EricTopol · 2026-10-09
- First Workshop on Agent Behavior at COLM 2026 Set for Oct 9 in San Francisco — _Hao_Zhu · 2026-10-09
- Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control — deliprao · 2026-10-09
- Autorubric at COLM 2026: a unifying framework for rubric-based LLM evaluation — deliprao · 2026-10-09
- TraceExtract open-sourced: data engine for µ0 world model trained on video with zero action labels — RexDouglass · 2026-10-09
- 500 curated SWE tasks lift Qwen 27B by 11.3 points in 15 GRPO steps — ycombinator · 2026-10-09