CertJudge Reuses Evaluations With One Calibration
sanmikoyejo · x · 2026-07-12
The overview notes that LLM judges are currently used to grade AI-generated Lean 4 theorems, but the problem is "who judges the judge." Every time a judge, prompt, or threshold is changed, a new manual study is typically required. The goal of CertJudge is to use a single small-scale calibration that can be reused for all subsequent judge evaluations, achieving a correlation of roughly ρ=0.833.
Related event: CertJudge Targets Reusable Calibration for LLM Judges(3 posts)→
More from Research
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- Yann LeCun live at ECCV on World Models — Weak_Assistance_5261 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11