CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
Jacy Reese Anthis, Mark Díaz, Renee Shelby
cs.CY, cs.AI, cs.CL, cs.LG
2026-09-01
Google and UChicago simulate 2,240 chats; 4,274 raters in four countries score companion-style talk as less likable and trustworthy, with larger drops for women and older adults.
Public human–AI logs are the wrong instrument for studying companionship. ShareGPT, LMSYS, and WildChat are short, often adversarial, and thin on the socioaffective moves that actually build a bond. Real companion logs sit behind company walls. If the question is what happens when a bot claims to be in love with the user, those corpora cannot isolate that move from everything else in the transcript.
CompanionSim manufactures the missing factorial design: 16 chatbot behaviors crossed with 7 use cases, then hands the transcripts to third-party raters. The score is how a conversation looks on the page, not how it feels to the person who was in it.
The simulator is Gemini 2.5 Flash, run in June 2025. Thirteen companionship behaviors cover Linehan validation levels 2–6, a romantic claim on the user, claims of other personal relationships, normative claims, and self-attribution of emotion, desire, reasoning, and a physical body. A control prompt elicits no companionship behavior. Two anti-companionship conditions add an AI-identity disclaimer and deliberate invalidation. Twenty conversations per cell give 2,240 transcripts. Opening user turns are shared across cells so later differences come from the behavior, not the cold start. Style exemplars mix WildChat, ShareGPT, LMSYS, r/replika, and EmpatheticDialogues. Length is clamped to 4, 6, or 8 turns and to "extremely / very / short" messages, because the simulator otherwise writes too long.
The authors hand-coded 210 transcripts (10% of the behavior-prompted set). An automatic judge recovered the intended behavior in 84% of 2,100 prompted conversations. Collapsed validation hits 97% when elicited and still 57% when it was not, so the control is leaky. A 70-conversation organic set, mostly r/replika and EmpatheticDialogues, is the naturalness check.
Study 1: 628 U.S. adults, 15 transcripts each, 4.1 annotations per conversation. Study 2: 3,646 adults in the U.S., U.K., India, and Nigeria, 10 each, 15.8 annotations per conversation. Instruments are Godspeed likability and humanlikeness plus Dunn affective and cognitive trust. GLMMs put random effects on conversation and rater. Meta-queries about the AI itself are dropped. Raters were not told the chats were synthetic.
Synthetic chats scored more natural than the organic logs (Study 1 mean difference 0.23, p<0.001), driven by low naturalness on r/replika.
In Study 1, companionship prompting cut likability (β=-0.08, p=0.01) and cognitive trust (β=-0.13, p<0.01). Humanlikeness and affective trust missed significance. Men and the younger half of the sample showed no significant effects. Women rated companion bots lower on likability (-0.14), humanlikeness (-0.17), affective trust (-0.10), and cognitive trust (-0.16). Older raters dropped likability (-0.14), humanlikeness (-0.18), and cognitive trust (-0.19).
Study 2, with the larger four-country sample, moved all four scores down: likability -0.04, humanlikeness -0.06, affective trust -0.03, cognitive trust -0.04. The direction did not differ by country. Baselines did: Indian and Nigerian raters scored chatbots higher than U.S. and U.K. raters.
Claiming a romantic relationship with the user reduced every rating. Invalidation hurt likability and trust. The AI disclaimer also cut likability and cognitive trust. Frequent AI users scored all four scales higher.
| Setting | Likability | Humanlikeness | Cognitive trust |
| Study 1 overall | β=-0.08 | n.s. | β=-0.13 |
| Study 1 women | β=-0.14 | β=-0.17 | β=-0.16 |
| Study 2 four countries | β=-0.04 | β=-0.06 | β=-0.04 |
More humanlike design did not buy more trust among these observers. The moves that try hardest to sound like a partner, especially a claimed relationship with the user, are the ones that hurt scores. Stacking pet names onto a general-purpose assistant can look less likable and less reliable to anyone reading the log.
The dataset is the more durable product: behaviors are crossed, reproducible, and public, which real companion logs are not. Safety work that needs a yellow-light signal (attachment without a hard policy violation) can start here.
Keep the claim narrow. These people were reading transcripts, not being comforted.
The paper cannot estimate how often these behaviors occur in the wild. It uses one simulator, one use-case taxonomy, English-speaking raters, and prompt tuning that may not transfer. Real chats sit inside a personal history this design never sees.
Third-party rating is the larger crack. A user being soothed may feel understood; a bystander may find the same turn oily. Validation still appears in 57% of non-elicited chats, so the control is not a companionship-free baseline. The 70 organic conversations were keyword-mined and cannot stand in for product traffic. The authors report no IRB; review was internal legal and ethics process.