EMNLP paper: LLM-simulated conversations are far less inconsistent than real humans
peizNLP · x · 2026-09-12
A paper by Ryo Kamoi et al., accepted to EMNLP 2026, questions whether LLMs can truly simulate human group behavior. The authors argue simulated conversations should reproduce inconsistent and uncollaborative behaviors (misunderstandings, interruptions) at human-like frequencies.
They introduce CoCoEval, an evaluation framework with:
- Turn-level detection of 10 types of inconsistent/uncollaborative behaviors
- A benchmark covering professional scenarios with collaboration and conflict
Comparing human conversations with simulations from GPT-4.1, GPT-5.1, and Claude Opus 4, they find that under vanilla prompting LLM-simulated dialogues show far fewer such behaviors than real humans — a caution for using LLMs as replacements for human studies in social science.
More from AGI Musings
- Bill Gates predicts major white-collar disruption within 2 years as AI reliability improves — rohanpaul_ai · 2026-09-12
- Polymarket odds hit 71% that an AI lab announces a Millennium Prize solution this year — Polymarket · 2026-09-12
- Oncologist Vinay Prasad: AI doom talk and cancer-cure promises are both marketing — VPrasadMDMPH · 2026-09-12
- Ex-Anthropic researcher's AI self-replication warning mocked: where are the 10,000 GB300 rigs? — basedjensen · 2026-09-12
- Fields Medalist Deng Yu says he'll quit math and write romance novels if AI solves all of it — Polymarket · 2026-09-12
- Manifold prices AI wiping out humanity by 2030 at just 5%, sparking a debate on doom markets — j_foerst · 2026-09-12