Five exercises find four defects in the author's own LLM-judge eval system

alexpran · reddit · 2026-09-15

Applying Dan Luu's 'what's wrong with this benchmark?' method to his own LLM-judge evals, the author surfaces four defects: non-unanimous rates swing from 2/21 to 6/21 across 16 identical replicate runs; a 17→144 case suite breaks the unchanged fail-threshold rule; precision computed on half the pipeline for two weeks looked plausible the whole time; and two sensible prompt edits dropped recall from 0.80 to 0.50 or narrowed the rule they meant to widen. Answers and run files in the first comment.

Original post →

More from coding & agent

coding & agent channel →