Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged
maverick_man1111 · reddit · 2026-10-06
The author spent a month reading the actual scoring code (not docs) of 7 tools used to grade Claude Code skills and LLM apps — the eval harness in the agent-skills repo, NVIDIA's SkillEvaluator, plus MLflow, LangSmith, DSPy, DeepEval, and Harbor — and found 13 cases where the reported pass/fail or score isn't backed by what the tool actually checked. Every case was reproducible, with writeups sent to maintainers.
Highlights:
- The judge model (usually Claude) replies "I can't score this" and the tool converts that into a score anyway.
- In MLflow, a gate requiring the new model to beat the old one by 10% breaks on negative scores (R²): the division flips, so the worse model passes and the better one fails — the author's fix was merged.
- Scores saved against the wrong run; leftover state from the previous run read as the current one.
All 13 issues were filed upstream, fixes sent for 12, with 6 merged so far. Three of the bugs were in the author's own tool Driftproof, an open-source Claude Code plugin for checking whether a skill still helps after a model update (v0.14.0 just released).
Key takeaway: running more evals doesn't help if the code checks the wrong thing — 100 runs just make you more confident in the wrong answer. Full writeup on Zenodo (DOI: 10.5281/zenodo.23050796).
More from coding & agent
- t3 code spawns agent types to review PRs with Alchemy's 20s deploys — samgoodwin89 · 2026-10-06
- Hugging Face adds an RL Environments filter to the Hub, supporting 4 frameworks — lmoroney · 2026-10-06
- Karpathy's arrival makes Anthropic's coders better, says Kuprel — Kuprel · 2026-10-06
- 'DOCX Will Be Replaced by Markdown Within 5 Years' Sparks Pushback from Formats Veteran — bytebot · 2026-10-06
- A browser tab refresh bug kept a server at 100% CPU for 4 days—load scaled with tabs squared — Ok_Negotiation_2587 · 2026-10-06
- Loop-engineering: 8 unattended agent loop patterns to run your repo while you sleep — JafarNajafov · 2026-10-06