NeurIPS paper proposes CAPA to show similar models may weaken AI oversight
dhadfieldmenell · x · 2026-07-25
The post points to a NeurIPS paper titled Great Models Think Alike and this Undermines AI Oversight. The authors propose CAPA, a chance-adjusted probabilistic agreement metric for model similarity based on overlap in mistakes. They report three main findings: LLM-as-a-judge scores are biased toward models more similar to the judge, weak-to-strong generalization works better when the weak supervisor and strong student are more different, and model errors are becoming more correlated as capabilities increase, which may weaken AI oversight.
More from Research
- An essay links compression and intelligence to mark Ray Solomonoff’s 100th birthday — ryangr · 2026-07-25
- 10 agent eval patterns every AI engineer should know, from golden sets to trajectory scoring — Roger_M_Taylor · 2026-07-25
- A copy-paste SOP aims to make LLM analysis more reliable — Local-Reading-1624 · 2026-07-25
- HUG uses 1M egocentric frames to train zero-shot robot grasping — chris_j_paxton · 2026-07-25
- A model screenshot admits it anthropomorphized itself, turning a technical caveat into a joke — sebkrier · 2026-07-25
- What data are labs using to train rumored 10T-parameter models? — Ill_Fisherman8352 · 2026-07-25