CertJudge Reuses Evaluations With One Calibration
sanmikoyejo · x · 2026-07-12
The overview notes that LLM judges are currently used to grade AI-generated Lean 4 theorems, but the problem is "who judges the judge." Every time a judge, prompt, or threshold is changed, a new manual study is typically required. The goal of CertJudge is to use a single small-scale calibration that can be reused for all subsequent judge evaluations, achieving a correlation of roughly ρ=0.833.
Related event: CertJudge Targets Reusable Calibration for LLM Judges(3 posts)→
More from Research
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- OpenAI and Apollo Research introduce Contrastive SDF to measure reward-seeking — OpenAI · 2026-07-22
- NVIDIA says to tune the harness before tuning the model with LangChain — NVIDIAAI · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22
- NVIDIA shows 22 SIGGRAPH papers and Omniverse tools for robot simulation — facontidavide · 2026-07-22
- Building a Knowledge Graph Without a Graph DB: 1000x Cheaper Than GraphRAG — TheRedfather · 2026-07-22