Certifying LLM Judges via Behavioral Checks

sanmikoyejo · x · 2026-07-12

The core idea is to stop constantly re-annotating, and instead "certify" judges using automated, falsifiable behavioral checks. The paper lists four categories of checks:\n\n1. P1 Identity\n2. P2 Bug monotonicity\n3. P3 Spec monotonicity\n4. P4 Stability\n\nBy applying controlled perturbations to arbitrary Lean files, this method verifies whether a judge is reliable without the need for gold-standard labels.

Original post →

More from Research

Research channel →