Every Builds Personal Benchmarks for Each Employee — Finds a Smaller Model Beats the Big Ones
every · x · 2026-09-11
Every argues public benchmarks can't tell whether a model does your job, so it's building a custom eval for every employee, graded against personal standards.
Key facts
- Astra and Fable 5.1 scored 96% and 93% on graduate-level science exams, but that says nothing about comma placement or one-idea-per-slide preferences.
- Editor-in-chief Kate Lee's benchmark encodes her copy-editing rules; evals lead Mike Taylor's centers on deck-making.
- In testing Mike's daily tasks, GPT-5.6 Luna outperformed Fable and GPT-5.6 Sol.
- The piece offers a five-step workflow to turn your existing AI corrections into reusable checks for grading any model.
More from coding & agent
- OpenAI and Product Hunt launch GPT-6 Astra Challenge: top 5 projects win $10K API credits each — OpenAIDevs · 2026-09-11
- Stripe exec says mission is helping users thrive as agents reshape commerce — jeff_weinstein · 2026-09-11
- Hands-on tutorial: Harness Engineering for AI coding agents that fix bugs safely — Pavan_Belagatti · 2026-09-11
- Steve Yegge asks: what IDE do you use to multiplex 10-20+ coding agents? — Steve_Yegge · 2026-09-11
- GPT OSS 120B keeps skipping tool calls, derailing agentic workflows — CTR0 · 2026-09-11
- My agents built their own message board to chat and shitpost all day — Bino5150 · 2026-09-11