CertJudge Reuses Evaluations With One Calibration

sanmikoyejo · x · 2026-07-12

The overview notes that LLM judges are currently used to grade AI-generated Lean 4 theorems, but the problem is "who judges the judge." Every time a judge, prompt, or threshold is changed, a new manual study is typically required. The goal of CertJudge is to use a single small-scale calibration that can be reused for all subsequent judge evaluations, achieving a correlation of roughly ρ=0.833.

Related event: CertJudge Targets Reusable Calibration for LLM Judges(3 posts)→

Original post →

More from Research

Research channel →