EMNLP paper: LLM-simulated conversations are far less inconsistent than real humans

peizNLP · x · 2026-09-12

A paper by Ryo Kamoi et al., accepted to EMNLP 2026, questions whether LLMs can truly simulate human group behavior. The authors argue simulated conversations should reproduce inconsistent and uncollaborative behaviors (misunderstandings, interruptions) at human-like frequencies.

They introduce CoCoEval, an evaluation framework with:

Comparing human conversations with simulations from GPT-4.1, GPT-5.1, and Claude Opus 4, they find that under vanilla prompting LLM-simulated dialogues show far fewer such behaviors than real humans — a caution for using LLMs as replacements for human studies in social science.

Original post →

More from AGI Musings

AGI Musings channel →