Evals aren't unit tests: scores fluctuate, so '85%' alone means nothing

emeka_boris · x · 2026-10-05

Developer chiziaruhoma calls out a common misconception: treating AI evals like unit tests. A unit test passes or fails deterministically, but an eval yields a score that shifts on every run — so "85%" in isolation is meaningless. What matters is across how many runs, which prompts, and which model settings. A useful reminder for agent/model eval engineering to account for the statistical nature of evals rather than applying binary-test thinking.

Original post →

More from coding & agent

coding & agent channel →