BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest
moschles · reddit · 2026-10-09
A 2024 ICML paper (BABA-is-AI, MIT + Virginia Tech) showed GPT-4o and Gemini-1.5 models "fail dramatically" when generalization requires manipulating and combining game rules. The poster asks whether tera-parameter agentic swarms that now ace ARC-AGI-3 and FrontierMath tier 4 can solve these small key-door puzzles — and if not, the paper's importance has only compounded. They suggest relaying it to Francois Chollet and the ARC Foundation as a candidate benchmark for ARC-AGI-4. Code and paper (arXiv 2407.13729) are public.
Related event: Two-Year-Old BABA-is-AI Paper Still Stumps SOTA Models(2 posts)→
More from AGI Musings
- Founder Sparks Backlash Claiming Anthropic Is Riskier to Humanity Than Superintelligence — dbasch · 2026-10-09
- Creator fires back at AI-art critics who ignore the craft behind quality AI work — Uncanny_Harry · 2026-10-09
- Independent researcher's phase-transition AI acceleration model hits all 6 pre-registered prediction windows — sadeyeprophet · 2026-10-09
- Curated reading list captures math community's debate on AI's impact on mathematics — StefanoGogioso · 2026-10-09
- Yudkowsky: AGI spend by 'end of the world' might not exceed one TSMC fab — panickssery · 2026-10-09
- VC thesis: everything reduces to 8 primitives and AI can 10-100x them all — signulll · 2026-10-09