Your LLM judge drifts whenever the hosted model silently updates
paratha27 · reddit · 2026-08-27
A practitioner shares his LLM-as-judge calibration workflow: the judge is used only for the one thing exact match can't check — whether a generated summary is faithful to its source — and he calibrates it against a hand-labelled set, reporting the agreement rate next to every score.
The catch: the judge runs on a hosted model that updates without warning, so the measuring instrument drifts on the same schedule as the thing being measured. He asks the community how large a calibration set needs to be before a drop in agreement means something rather than noise, and whether anyone has actually caught a judge going bad this way.
More from coding & agent
- LangChain Managed Deep Agents Support Environment Baking at Deploy — LangChain · 2026-08-27
- Why AI Agents Actually Need Memory? A Deep Dive into Technical Necessity — _jaydeepkarale · 2026-08-27
- ARK launches SDK to intercept bad tool decisions and enforce policies at runtime — Aromatic-Ad-6711 · 2026-08-27
- Agent Workforce Performance Depends on Setup — nikvassev · 2026-08-27
- Preventing Agents from Rewriting Contracts: A Three-Layer Architecture — haandol-_- · 2026-08-27
- Agent Autonomy Demo: Domain Bought Automatically Last Night — billyjhowell · 2026-08-27