AI agent panels agreed 98.7% of the time on a logic puzzle — and were wrong in 74% of cases
alex_verem · x · 2026-10-07
A Waseda University researcher replayed 100 real group discussions of a card logic puzzle with AI twins of each participant, stress-testing the growing practice of using AI personas as stand-ins in focus groups and survey panels.
- AI agent groups agreed on an answer 98.7% of the time, but in 74% of those groups they agreed on the wrong answer. Real groups reached consensus only 51.1% of the time.
- With thinking mode on, the model behind the twins (DeepSeek V4 Flash) reached consensus 95.6% of the time vs 85.2% with it off.
- Renaming the cards (same logic, invalidating the textbook answer) collapsed correct answers from 84% to 24.7% of groups — surface pattern matching, not reasoning. Qwen3-14B over-agreed on the same groups.
- About 20% of humans never wrote a message; only 0.2% of agents stayed silent. On an opinion question, agents agreed less than humans, so the error runs both ways.
Takeaway: unanimous agreement from a simulated AI panel tells you nothing about truth. Verify answers yourself.
Related event: Waseda Study Finds AI Focus Groups Agree Often — But Often Wrong Together(2 posts)→
More from AGI Musings
- 1,200 people reproduce 2,226 ICML papers with agents; AI reviewers diverge from humans — evijit · 2026-10-07
- Mathematician Daniel Litt: the math profession will change enormously, but nobody knows how — littmath · 2026-10-07
- 'The only moat left is caring about your project past two days' — danshipper · 2026-10-07
- Gary Marcus Asks: Is a 'Nice Tool' Worth a 10% Risk of Catastrophe? — GaryMarcus · 2026-10-07
- Gary Marcus: AI extinction risk near zero, but catastrophe and dystopia risks loom — GaryMarcus · 2026-10-07
- PhD student laments shrinking research horizons as IceCube's 38-year path wins the Nobel — DJiafei · 2026-10-07