Same AI agent scores 97% or 39% depending on how you define success

hugobowne · x · 2026-09-21

Hugo Bowne-Anderson illustrates how eval definitions change everything, using an Anthropic agent example:

Same agent, three very different numbers. If you can check multiple proposed fixes and keep the passing one, success-within-a-few-attempts is useful; if customers rely on the agent every time, consistency is what matters.

He's teaching a free Lightning Lesson, "AI Agent Evals: Test What Matters for Your Agent" (Sep 21/22), covering turning failures into eval cases, choosing between code checks, LLM judges and human review, and verifying that changes to your agent actually help.

Related event: Agent eval success rates hinge on how you define success(2 posts)→

Original post →

More from coding & agent

coding & agent channel →