Yoav Goldberg: Any Scorable Benchmark Will Eventually Be Brute-Forced
NLP researcher Yoav Goldberg argues that any reliably scorable benchmark will eventually be solved with enough brute-force training, though the resulting capabilities may not transfer. Citing an arXiv paper, Gavin Leech noted 78% of CodeForces problems have semantic duplicates in training corpora, further questioning benchmark validity.
2026-09-04 ~ 2026-09-04 · 3 related posts
- Yoav Goldberg: Every popular benchmark will be brute-forced away — so how do we measure real progress? — yoavgo · 2026-09-04
- Study finds semantic dupes for 78% of CodeForces problems in training data — gleech · 2026-09-04
1 near-duplicate retellings: herbiebradley