MMTEB Leaderboard Suspected of Training Set Contamination

antoine_chaffin · x · 2026-07-15

This reply adds an important observation about the MMTEB leaderboard: among its two official tasks, many models do not perform well.

The author notes that the few models that "look good" often directly used the training sets of those datasets, rather than demonstrating true generalization. This means leaderboard results may be significantly affected by data leakage / training set contamination; in other words, some top rankings do not indicate stronger models on broader tasks.

Related event: MMTEB Leaderboard Faces Training Data Contamination Controversy(3 posts)→

Original post →

More from Research

Research channel →