Your LLM judge drifts whenever the hosted model silently updates

paratha27 · reddit · 2026-08-27

A practitioner shares his LLM-as-judge calibration workflow: the judge is used only for the one thing exact match can't check — whether a generated summary is faithful to its source — and he calibrates it against a hand-labelled set, reporting the agreement rate next to every score.

The catch: the judge runs on a hosted model that updates without warning, so the measuring instrument drifts on the same schedule as the thing being measured. He asks the community how large a calibration set needs to be before a drop in agreement means something rather than noise, and whether anyone has actually caught a judge going bad this way.

Original post →

More from coding & agent

coding & agent channel →