Kitaru ships open-source skill to calibrate LLM judge scores against human labels

strickvl · x · 2026-09-24

Judge scores are probabilities, not measured accuracy — strickvl warns you must calibrate them against your own labels before gating on them. Kitaru ships a validate-evaluator skill that walks you through checking one evaluator against criterion-specific verdicts, inspecting disagreements, and assessing a frozen config on held-out cases.

He also verified the judge isn't just number-matching: a '30 days' from the policy tool passes, while promising a customer 30 days fails — even though a 30 appears in the tool results (it's the return window, not the refund time).

Original post →

More from coding & agent

coding & agent channel →