How often should you run evals? Hamel Husain lays out a 3-factor tradeoff framework
HamelHusain · x · 2026-10-06
AI evals expert Hamel Husain (with Shreya Shankar) answers "how often should I run my evals?" with three tradeoff dimensions:
- Cost to run and maintain: the more expensive the eval (e.g., LLM-as-a-judge vs. a unit test), the greater the benefit needed to justify frequent runs.
- Saturation: if the eval passes all examples, it gives no new information — retire it or run it less often; but first try making it harder rather than dropping it.
- Business value of catching the error: critical errors justify frequent runs despite cost. Red-team your app first to prove the error can trigger at least once before building the eval.
There are no bright-line rules. Concrete examples: a saturated answer-quality judge with GPT-6 Astra on max reasoning (medium failure cost) should be retired or run biweekly; a cheap code assertion checking a contact saved correctly can run on every change.
More from coding & agent
- He had Claude build Pokémon-style Korean-learning games, free to play and open source — EstanislaoStan · 2026-10-06
- After 2 years, a dev unveils Aether: a local AI 'OS' with structured memory and FailureMesh recovery — Budget_One_8784 · 2026-10-06
- Why LLMs never pick Java: dev points to under-sampling in training data — HanchungLee · 2026-10-06
- 12 practical Grok Bot use cases: from post-meeting tasks to inbox triage — brandon_galang · 2026-10-06
- StackOverflow survey: 65% of developers use coding agents, yet six in ten avoid AI — pchandrasekar · 2026-10-06
- teenytiny.computer launches cloud Linux VMs built for humans and agents to share — pritisinghhhh · 2026-10-06