NeurIPS paper proposes CAPA to show similar models may weaken AI oversight

dhadfieldmenell · x · 2026-07-25

The post points to a NeurIPS paper titled Great Models Think Alike and this Undermines AI Oversight. The authors propose CAPA, a chance-adjusted probabilistic agreement metric for model similarity based on overlap in mistakes. They report three main findings: LLM-as-a-judge scores are biased toward models more similar to the judge, weak-to-strong generalization works better when the weak supervisor and strong student are more different, and model errors are becoming more correlated as capabilities increase, which may weaken AI oversight.

Original post →

More from Research

Research channel →