MMTEB Leaderboard Suspected of Training Set Contamination
antoine_chaffin · x · 2026-07-15
This reply adds an important observation about the MMTEB leaderboard: among its two official tasks, many models do not perform well.
The author notes that the few models that "look good" often directly used the training sets of those datasets, rather than demonstrating true generalization. This means leaderboard results may be significantly affected by data leakage / training set contamination; in other words, some top rankings do not indicate stronger models on broader tasks.
Related event: MMTEB Leaderboard Faces Training Data Contamination Controversy(3 posts)→
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- Structural ensembles, not single predictions, drive robust TCR:pMHC generalization — quaidmorris · 2026-07-22
- enFoldX turns AlphaFold3 ensemble noise into a TCR–peptide–MHC predictor — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22