Stop Trusting MMLU-Pro: Every's Evals Lead Says Build Your Own AI Benchmark
danshipper · x · 2026-09-22
Mike Taylor, head of evals at Every, argues that MMLU-Pro-style benchmarks test trivia (like the cranial capacity of Homo erectus) rather than what matters: whether a model helps with your work. Drawing on his consulting practice—getting models to write copy, build dashboards, and assemble slide decks—he argues for building a personal benchmark: scoring new models on your real tasks instead of public leaderboards, since no one hires a VP based on SAT scores.
More from Research
- September 2026 robotics: Helix 2.5 in 30 homes, Digit 5 lifts 22.7 kg, OpenAI eyes humanoid — TheTuringPost · 2026-09-22
- New paper: what empirical evidence says about work, wellbeing, and AI futures — scychan_brains · 2026-09-22
- Multi-agent paper authors: periodically flushing context and keeping a summary doc helps a lot — DimitrisPapail · 2026-09-22
- Why multi-agent wins: Papail suggests entropy injection across API calls helps — DimitrisPapail · 2026-09-22
- Multi-agent vs serial debate: best agent needs 10-100x tokens to match team — generatorman_ai · 2026-09-22
- Nature Computational Science review maps four roles of LLMs as human proxies — iyadrahwan · 2026-09-22