NeurIPS paper proposes CAPA to show similar models may weaken AI oversight
dhadfieldmenell · x · 2026-07-25
The post points to a NeurIPS paper titled Great Models Think Alike and this Undermines AI Oversight. The authors propose CAPA, a chance-adjusted probabilistic agreement metric for model similarity based on overlap in mistakes. They report three main findings: LLM-as-a-judge scores are biased toward models more similar to the judge, weak-to-strong generalization works better when the weak supervisor and strong student are more different, and model errors are becoming more correlated as capabilities increase, which may weaken AI oversight.
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11