Every builds personal benchmarks for every employee: 'Your taste is test data'
danshipper · x · 2026-09-19
Every CEO Dan Shipper argues public benchmarks can't tell whether a model knows your deck needs one idea per slide — so Every is building a personal benchmark for every employee, grading models against your own standard of good.
- Astra and Fable 5.1 scored 96% and 93% on graduate-level science questions, but that says little about real work value.
- Editor in chief Kate Lee's benchmark encodes her copy-editing judgment; evals head Mike Taylor's deck benchmark convinced him a smaller model could handle much of his daily work.
- The piece offers a five-step workflow to turn the corrections you already give AI into reusable evals.
Core thesis: "Turn your corrections into evals. Your taste is test data."
More from coding & agent
- AI agent books a United flight change on its own as 'do it for me' bar rises — armand_ruiz · 2026-09-19
- Clinical agent run costs $0.0047: 99 calls, 143K tokens with Jev in a triage workflow — MaziyarPanahi · 2026-09-19
- Alchemy IaC adds secret manager support for Doppler, Infisical and more — samgoodwin89 · 2026-09-19
- Swarms to ship MCP Deployer in v16, turning any agent into an MCP server in 5 lines — KyeGomezB · 2026-09-19
- Letting agents deploy your app then browse it like real users: super high-level fuzzing — lucasmeijer · 2026-09-19
- Prompting rule: ban fallbacks so AI models can't cheat around core work — Daniel_Farinax · 2026-09-19