Every Builds Internal Platform for Personal Benchmarks Based on Real Work, Not Leaderboards
danshipper · x · 2026-09-12
Every's Dan Shipper argues benchmark scores say little about how a model performs on your actual work, which is why for three years the team has done "vibe checks": hands-on long-form reviews based on real work. Now they're doubling down quantitatively — @hammermt and @nityeshaga built an internal platform so everyone can create personal benchmarks from their day-to-day work, turning vibes into measurable checks.
Related event: Every Ditches Leaderboards, Builds Personal Benchmarks for Real Work(2 posts)→
More from Models
- Zed founder shows one link task eating ~50% of Claude subscription quota — zeeg · 2026-09-12
- Demo claims precise camera control with GPT-6 Astra (unverified) — round · 2026-09-12
- Burkov: Visual Reasoning Still Trails Luna and Gemini Flash — burkov · 2026-09-12
- Burkov on RL Whack-a-Mole: Fix One Failure, Another Pops Up — burkov · 2026-09-12
- GPT-Live-1 listens and speaks at once, shifting AI from prompts to live collaboration — WirelessLife · 2026-09-12
- Evals are increasingly unreadable: overly strict hidden tests mislabel good answers — trq212 · 2026-09-12