Yoav Goldberg: Every popular benchmark will be brute-forced away — so how do we measure real progress?

yoavgo · x · 2026-09-04

NLP researcher Yoav Goldberg offers a sober take on measuring AI progress: any task you can reliably score will be solved if you throw enough money ("an obscene amount") at brute-forcing its training; every challenging and popular benchmark will meet that bar. The catch — abilities won't necessarily transfer beyond the benchmark. His closing question: how can we actually measure progress?

Original post →

More from AGI Musings

AGI Musings channel →