Frontier Model Council Study: Models Fail Together
soumitrashukla9 · x · 2026-07-14
The authors re-ran frontier model experiments using the same initial benchmarks, including GPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, Grok 4.5. They developed an "error correlation" metric to measure the probability of two models answering incorrectly simultaneously.
They found:
- No council structure surpasses its strongest member on any benchmark.
- Models rarely disagree; across all 329 questions, there was never a case where "only one model got it right."
- When they answer incorrectly, they tend to fail together.
The author believes this supports the "consensus machines" argument: sharing training data, distillation targets, and RLHF habits causes models to converge on the same wrong conclusions. The text also notes that 6 pairs of vendors exhibited similar phenomena on MMLU-Pro Math.
More from Research
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- BlackboxNLP 2026 is recruiting extra reviewers after a high submission volume — hanjie_chen · 2026-07-22
- AWS shows self-distilled reasoning can preserve math and coding skills during SFT — AWS ML Blog · 2026-07-22
- UI2App shows screenshot fidelity still lags real interaction recovery — Grace Man Chen · 2026-07-22