Hamel Husain: refresh stale eval datasets with regular error analysis, retire passing ones
HamelHusain · x · 2026-09-23
In their AI Evals FAQ, Hamel Husain and Shreya Shankar answer what to do when your gold eval dataset goes stale:
- Evals naturally go stale as products and users change; use regular error analysis to find new problems and update examples/reference answers, with review cadence tied to how fast your product evolves
- Evals catch regressions like unit tests but cost far more to maintain — weigh cost vs signal value; if everything keeps passing, retire or run the eval less often
- Scores losing comparability with old runs is fine: evals exist to provide challenges you can hill-climb against
- For long-horizon progress tracking, lean on product metrics (churn, active users, revenue) alongside eval scores
More from coding & agent
- Jev API explodes at $0.042/M tokens: a hands-on checklist from desktop agents to drone control — blaizedsouza · 2026-09-23
- The prompt to run before wiring an agent to another service — gethackteam · 2026-09-23
- DeepSeek V4.1 Flash inside Codex harness impresses: steerable reasoning, promising results — Small_Ninja2344 · 2026-09-23
- Vercel CEO Backs px0, a Lightweight IDE Built for Reviewing Agent-Written Code — arpit_bhayani · 2026-09-23
- Full course: automating SEO and AEO with AI agents built on Opus 5.5 — eptwts · 2026-09-23
- Matt Pocock: Stop chasing model releases, improve your agent's harness instead — mattpocockuk · 2026-09-23