Why decontamination reports can't fix benchmark contamination — and what evaluators must do instead
NoahPersaud · reddit · 2026-09-20
After OpenAI retired SWE-bench Verified (every frontier model could reproduce reference fixes verbatim; scores rose just 6 points in 6 months), the author argues decontamination reports are structurally incapable: labs check their own undisclosed corpora, corpora can't be published for litigation reasons, and n-gram matching misses paraphrases, forum walkthroughs, GitHub solutions, and synthetic data. Commitments and PSI prove only what labs declared, and proof-of-training schemes have been spoofed. His fix: evaluators control the test — no labels to submitters, offline evaluation, builds from named commits, test data generated after submission freeze, results only count if reproduced. He built a small version and honestly lists its remaining gaps.
More from Research
- Have we seen an acceleration in discoveries? Cyber spikes, math rises, algorithms flat — soumitrashukla9 · 2026-09-20
- TovanaEngine: a local world model trained on 50K+ real SWE-bench coding-agent runs — Decent-Ad9950 · 2026-09-20
- After Chess and Go: Can Any AI Engine Actually Beat Humans at Scrabble? — zuilserip · 2026-09-20
- Textbook author: 99.9% accuracy can mean zero scientific discoveries — bravo_abad · 2026-09-20
- rasbt: Jev's Secret Sauce Is Data, Not the Algorithm — Laya Rival Falls Far Short — RichmanRonald · 2026-09-20
- Sebastian Raschka open-sources an end-to-end 'AI text detector from scratch' project — rasbt · 2026-09-20