Human Consistency Cannot Validate Judges

sanmikoyejo · x · 2026-07-12

The thread points out that using human consistency to validate judges is not scalable: any configuration change invalidates previous rounds of research. The Lean kernel is also unhelpful here because it only checks whether a proof passes; it cannot determine if the proposed theorem is correct, too broad, or under-constrained. Even if the code typechecks, it could still have semantic issues.

Original post →

More from Research

Research channel →