Yoav Goldberg: every popular benchmark will be brute-forced, without real transfer
herbiebradley · x · 2026-09-04
Researcher Yoav Goldberg lays out a skeptical framework for measuring AI progress: any task with reliable scoring can be solved by brute-forcing training with enough money, every challenging popular benchmark meets that bar, and abilities won't necessarily transfer — so how do we measure real progress? Replies raise human-in-the-loop / centaur evals as one possible answer.
Related event: Yoav Goldberg: Any Scoreable Benchmark Falls to Brute-Force Training(2 posts)→
More from AGI Musings
- Paradigm 3: low-quality RL environments may explain reward hacking; EBR-bench shows humans beat AIs — gleech · 2026-09-04
- Zero failure rate on alignment evals is a red flag, warn safety researchers — connoraxiotes · 2026-09-04
- Skeptical take: OpenAI can't train large models, pivots to RL and inference — teortaxesTex · 2026-09-04
- Safety researcher invokes professional standards to question Altman's safety claims — davidmanheim · 2026-09-04
- 100% AI-powered media reportedly beats journalists to an OpenAI scoop — emmanuelvivier · 2026-09-04
- Mathematicians Grumble as AI Cracks Conjectures 'The Wrong Way' — avt_im · 2026-09-04