Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged

maverick_man1111 · reddit · 2026-10-06

The author spent a month reading the actual scoring code (not docs) of 7 tools used to grade Claude Code skills and LLM apps — the eval harness in the agent-skills repo, NVIDIA's SkillEvaluator, plus MLflow, LangSmith, DSPy, DeepEval, and Harbor — and found 13 cases where the reported pass/fail or score isn't backed by what the tool actually checked. Every case was reproducible, with writeups sent to maintainers.

Highlights:

All 13 issues were filed upstream, fixes sent for 12, with 6 merged so far. Three of the bugs were in the author's own tool Driftproof, an open-source Claude Code plugin for checking whether a skill still helps after a model update (v0.14.0 just released).

Key takeaway: running more evals doesn't help if the code checks the wrong thing — 100 runs just make you more confident in the wrong answer. Full writeup on Zenodo (DOI: 10.5281/zenodo.23050796).

Original post →

More from coding & agent

coding & agent channel →