Why decontamination reports can't fix benchmark contamination — and what evaluators must do instead

NoahPersaud · reddit · 2026-09-20

After OpenAI retired SWE-bench Verified (every frontier model could reproduce reference fixes verbatim; scores rose just 6 points in 6 months), the author argues decontamination reports are structurally incapable: labs check their own undisclosed corpora, corpora can't be published for litigation reasons, and n-gram matching misses paraphrases, forum walkthroughs, GitHub solutions, and synthetic data. Commitments and PSI prove only what labs declared, and proof-of-training schemes have been spoofed. His fix: evaluators control the test — no labels to submitters, offline evaluation, builds from named commits, test data generated after submission freeze, results only count if reproduced. He built a small version and honestly lists its remaining gaps.

Original post →

More from Research

Research channel →