Two-Year-Old BABA-is-AI Paper Still Stumps SOTA Models
A 2024 ICML paper by MIT and Virginia Tech researchers found that SOTA multimodal models like GPT-4o and Gemini-1.5 fail badly on BABA puzzles requiring certain manipulations. Reddit users are now calling for retesting newer models against the benchmark.
2026-10-09 ~ 2026-10-09 · 2 related posts
- BABA-is-AI: 2024 ICML benchmark that broke SOTA LLMs deserves a 2026 retest — moschles · 2026-10-09
1 near-duplicate retellings: moschles