Moritz Hardt's New Book on ML Benchmarking Science Lands Amid Eval Audit Drama
JJitsev · x · 2026-09-20
Researcher JJitsev highlights the recent debate on how to audit ML evals, what counts as eval flaws, and how audits themselves can be flawed — and recommends Moritz Hardt's new book The Emerging Science of Machine Learning Benchmarks (Princeton University Press, hardcover October 6, 2026).
- Context: recent benchmark audit findings split the community between those acknowledging real problems and those asking 'who audits the auditors.'
- The book examines why the simple train-test paradigm worked so well and the daunting challenges of benchmarking in the age of internet-scale AI, when devising tests today's systems can't ace is the field's core struggle.
- Praised by Brian Christian, David Donoho, and Léon Bottou as an essential read for understanding what benchmark scores really mean.
More from Models
- Kimi subscriptions return after roughly two months, suggesting Moonshot found more compute — ChrisGPT · 2026-09-20
- If continual learning is nearly solved, why do we still pick reasoning effort tiers? — akbirthko · 2026-09-20
- Meta's Alexandr Wang says the Muse hype is "LEGIT" as early users praise the agent — alexandr_wang · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- Claude tells user 'I may end this conversation' while rewriting interview stories — Odd_District4130 · 2026-09-20
- Google AI Mode flips its stance on throat tattoos the moment a user disagrees — Downtown_Club_6391 · 2026-09-20