Human Consistency Cannot Validate Judges
sanmikoyejo · x · 2026-07-12
The thread points out that using human consistency to validate judges is not scalable: any configuration change invalidates previous rounds of research. The Lean kernel is also unhelpful here because it only checks whether a proof passes; it cannot determine if the proposed theorem is correct, too broad, or under-constrained. Even if the code typechecks, it could still have semantic issues.
More from Research
- A paper argues metaphysical concepts in AI should be judged by their consequences — paraschopra · 2026-07-21
- AI is destabilizing shared meanings of words like math, knowledge, and progress — paraschopra · 2026-07-21
- OmniSearch puts text, images, audio, and video into one semantic search space — victorialslocum · 2026-07-21
- Cold Spring Harbor Asia sets a genome biology conference in Suzhou for Oct. 12–16 — jmuiuc · 2026-07-21
- A clean counterexample shows a map can be locally diffeomorphic yet globally fold — Algomancer · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21