Frontier Model Council Study: Models Fail Together
soumitrashukla9 · x · 2026-07-14
The authors re-ran frontier model experiments using the same initial benchmarks, including GPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, Grok 4.5. They developed an "error correlation" metric to measure the probability of two models answering incorrectly simultaneously.
They found:
- No council structure surpasses its strongest member on any benchmark.
- Models rarely disagree; across all 329 questions, there was never a case where "only one model got it right."
- When they answer incorrectly, they tend to fail together.
The author believes this supports the "consensus machines" argument: sharing training data, distillation targets, and RLHF habits causes models to converge on the same wrong conclusions. The text also notes that 6 pairs of vendors exhibited similar phenomena on MMLU-Pro Math.
More from Research
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11
- AnyMatch accepted to ECCV 2026 with a NoPresenter design — ducha_aiki · 2026-09-11