Dev Slams Benchmark Contamination: Don't Train on BEIR and Evaluate on It
bo_wangbo · x · 2026-10-09
In the reranker debate, bowangbo pinpoints the core issue: BEIR itself is a great benchmark, but the common practice of training on BEIR, evaluating on BEIR, and then claiming the model is good amounts to benchmark contamination. His advice: build your own secret benchmark to measure true generalization instead of relying on contaminated public leaderboards.
More from Research
- PersistBench wins NeurIPS Spotlight, finds 4D foundation models lack visual memory — weichiuma · 2026-10-09
- StarkWare founder Eli Ben-Sasson: AI solved the Erdős Unit Distance problem, all bets are off — jamestagg · 2026-10-09
- Claude Science produces first complete ultraviolet map of the sky, ~10% deviation — The Decoder · 2026-10-09
- FreeMatching: generalizable dense correspondence matching beyond spatio-temporal priors — hkuhk · 2026-10-09
- MBZUAI's WorldGuide beats MiniMax-H3 on closed-loop procedural video world modeling — MBZUAI · 2026-10-09
- REMORY adds soft residual memory tokens to context compaction, hitting near full-context scores at 5.2% of input — Hanchen Xia · 2026-10-09