Dev Slams Benchmark Contamination: Don't Train on BEIR and Evaluate on It

bo_wangbo · x · 2026-10-09

In the reranker debate, bowangbo pinpoints the core issue: BEIR itself is a great benchmark, but the common practice of training on BEIR, evaluating on BEIR, and then claiming the model is good amounts to benchmark contamination. His advice: build your own secret benchmark to measure true generalization instead of relying on contaminated public leaderboards.

Related event: Reranker Debate: New Paper Lifts BEIR Scores Amid Benchmark Contamination Concerns(2 posts)→

Original post →

More from Research

Research channel →