Yoav Goldberg: Every popular benchmark will be brute-forced away — so how do we measure real progress?
yoavgo · x · 2026-09-04
NLP researcher Yoav Goldberg offers a sober take on measuring AI progress: any task you can reliably score will be solved if you throw enough money ("an obscene amount") at brute-forcing its training; every challenging and popular benchmark will meet that bar. The catch — abilities won't necessarily transfer beyond the benchmark. His closing question: how can we actually measure progress?
More from AGI Musings
- Forethought weighs a superintelligent "nightwatchman" aboard galactic colonization probes — willmacaskill · 2026-09-04
- AI BioDesign accelerator launches with UW Medicine and Fred Hutch to let AI design biology — AllThingsApx · 2026-09-04
- Uber now fights self-driving cars it once personified, and AI gains may stall — carlbfrey · 2026-09-04
- Conceding to a commenter: treating AI consciousness uncertainty like engineering safety margins — Passelume · 2026-09-04
- Move 37 vs Zelnick: why datasets are a floor for creativity, not a ceiling — r0ck3t23 · 2026-09-04
- "We already have AGI in software development": a narrow take on Astra — haider1 · 2026-09-04