Yoav Goldberg: every popular benchmark will be brute-forced, without real transfer

herbiebradley · x · 2026-09-04

Researcher Yoav Goldberg lays out a skeptical framework for measuring AI progress: any task with reliable scoring can be solved by brute-forcing training with enough money, every challenging popular benchmark meets that bar, and abilities won't necessarily transfer — so how do we measure real progress? Replies raise human-in-the-loop / centaur evals as one possible answer.

Related event: Yoav Goldberg: Any Scoreable Benchmark Falls to Brute-Force Training(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →