Replaying 100 human group discussions with LLM agents: 98.7% consensus, mostly on wrong answers
alex_verem · x · 2026-10-07
A Waseda University researcher's arXiv paper (2609.20543) replayed 100 real human Wason-card group discussions with belief-anchored LLM agent groups.
Key findings:
- Agent groups agreed 98.7% of the time on the logic puzzle — but in 74% of those groups they agreed on a wrong answer
- Human full-consensus rates range 24%-57% depending on scoring definitions, and about one fifth of human participants never posted, while agents almost always did
- Submit-based comparisons show 34.0-43.9 percentage-point consensus gaps (chat vs reasoning modes), and participation-matched comparisons converge within 0.5 points
- The gap persisted without early stopping and after removing the memorizable answer; reasoning-mode groups approached unanimous agreement, mostly on incorrect answers
Conclusion: simulated consensus does not track collective accuracy, and belief-anchored agent groups are biased estimators of real human deliberation — a warning for social-simulation research.
Related event: Waseda Study Finds AI Focus Groups Agree Often — But Often Wrong Together(2 posts)→
More from AGI Musings
- Pilot study: AI agents rank all ICML 2026 papers as area chairs, and their 'research taste' differs from humans — ShayneRedford · 2026-10-07
- Neuroscientists debate whether brain properties left out of computational models matter for consciousness — eschwitz · 2026-10-07
- 1,200 people reproduce 2,226 ICML papers with agents; AI reviewers diverge from humans — evijit · 2026-10-07
- Mathematician Daniel Litt: the math profession will change enormously, but nobody knows how — littmath · 2026-10-07
- 'The only moat left is caring about your project past two days' — danshipper · 2026-10-07
- Gary Marcus Asks: Is a 'Nice Tool' Worth a 10% Risk of Catastrophe? — GaryMarcus · 2026-10-07