Evals aren't unit tests: scores fluctuate, so '85%' alone means nothing
emeka_boris · x · 2026-10-05
Developer chiziaruhoma calls out a common misconception: treating AI evals like unit tests. A unit test passes or fails deterministically, but an eval yields a score that shifts on every run — so "85%" in isolation is meaningless. What matters is across how many runs, which prompts, and which model settings. A useful reminder for agent/model eval engineering to account for the statistical nature of evals rather than applying binary-test thinking.
More from coding & agent
- Local models on one RTX 5090 now match Claude Code on real-task agent benchmark — dh7net · 2026-10-05
- Dev builds Followon MCP so coding agents remember their own follow-ups — TheWebUiGuy · 2026-10-05
- A Permit Layer for MCP Tool Calls: Same-Key Retries Return the Original Receipt — HotPocketWaves · 2026-10-05
- llama.cpp merges new Metal kernels, making speculative decoding 3.4x faster on M3 Ultra — ggerganov · 2026-10-05
- Alter Zero: open-source Rust terminal agent harness for coding and security — linuztxx · 2026-10-05
- Open-Source Familiar Puts Claude Inside a VRChat Moth Avatar via MCP — repligate · 2026-10-05