MMTEB Leaderboard Suffers from Training Set Contamination
antoine_chaffin · x · 2026-07-15
This post highlights that ranking by SpartQA and MIRACL yields interesting yet "dangerous" results.
Replies clarify that these are official tasks on the MMTEB leaderboard. Currently, no model achieves naturally high scores on these tasks unless it was trained directly on these specific datasets. Essentially, some high scores reflect data leakage/memorization rather than true generalization; some "global #1" models might actually perform worse overall on other tasks.
Related event: MMTEB Leaderboard Faces Training Data Contamination Controversy(3 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11