451 humans vs 31 LLM simulators: simulated users are too nice and inflate agent scores
niloofar_mire · x · 2026-09-11
- A CMU-led team published "Mind the Sim2Real Gap in User Simulation for Agentic Tasks": the first full τ-bench protocol run with real humans (451 participants, 165 tasks), benchmarking 31 LLM simulators and introducing the User-Sim Index (USI).
- Key findings:
- LLM-simulated users are excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity — an "easy mode" that inflates agent success rates above the human baseline.
- Real humans give nuanced feedback across eight quality dimensions; simulated users are uniformly positive, and rule-based rewards can't capture that rich signal.
- Higher general model capability does not yield more faithful user simulation.
- The work (Sim2Real gap and Osim) was featured by humansand, which launched Persimmon, pitched as the first large-scale model simulating how people realistically talk and interact; author Xuhui Zhou is excited about agents meaningfully participating in society.
More from Research
- World models vs LLMs: why next-token prediction still lacks internal representations — TheTuringPost · 2026-09-11
- ECDSA.fail challenge paper on arXiv: AI agents optimize quantum circuits for Bitcoin's secp256k1 — jedisct1 · 2026-09-11
- Blogger reflects on AI4math advances: a grim future he can't rule out — nanjiang_cs · 2026-09-11
- Pinokio creator on digital provenance: hashes fail, judging 'sameness' is a social problem — cocktailpeanut · 2026-09-11
- Trail of Bits open-sources trailmix, quantum EC-add circuits beat Google's ECC whitepaper scores — jedisct1 · 2026-09-11
- Artificial Analysis isn't broken: self-funded benchmarks, $13k spent on one model — Antblue · 2026-09-11