Stop Trusting MMLU-Pro: Every's Evals Lead Says Build Your Own AI Benchmark

danshipper · x · 2026-09-22

Mike Taylor, head of evals at Every, argues that MMLU-Pro-style benchmarks test trivia (like the cranial capacity of Homo erectus) rather than what matters: whether a model helps with your work. Drawing on his consulting practice—getting models to write copy, build dashboards, and assemble slide decks—he argues for building a personal benchmark: scoring new models on your real tasks instead of public leaderboards, since no one hires a VP based on SAT scores.

Original post →

More from Research

Research channel →