Kitaru ships open-source skill to calibrate LLM judge scores against human labels
strickvl · x · 2026-09-24
Judge scores are probabilities, not measured accuracy — strickvl warns you must calibrate them against your own labels before gating on them. Kitaru ships a validate-evaluator skill that walks you through checking one evaluator against criterion-specific verdicts, inspecting disagreements, and assessing a frozen config on held-out cases.
He also verified the judge isn't just number-matching: a '30 days' from the policy tool passes, while promising a customer 30 days fails — even though a 30 appears in the tool results (it's the return window, not the refund time).
More from coding & agent
- "This Is Too Slow, Make It Faster" — Dev's One-Liner Actually Optimized His AI Code — DanielLockyer · 2026-09-24
- SwiftFairy launches: on-device Mac app that reviews your coding agent's Swift code — JordanMorgan10 · 2026-09-24
- Stripe's Link Wallet Tops 300M Consumers, Partners With Muse for Agentic Buying — jeff_weinstein · 2026-09-24
- Zuckerberg Put Cameras in His MMA Gym So an Agent Could Coach Him Between 1-Minute Rounds — victor_explore · 2026-09-24
- LangChain's Interrupt keynote ships LangSmith Engine v2, Deep Agents 0.8, fine-tuning and more — LangChain · 2026-09-24
- Runable launches Scheduled Tasks: agents that do the work, not just remind you — SimplyAnnisa · 2026-09-24