Year-long Character.AI study: sustained companionship tracks lower well-being via less in-person contact

Living with AI Companions: Sustained AI Companionship Predicts Lower Well-Being Through Lower Human Interaction

Yutong Zhang, Dora Zhao, Yixin Wang, Rebecca Anselmetti, Jeffrey T. Hancock, Robert Kraut, Diyi Yang

cs.HC

2026-09-07

A two-wave study of 1,182 Character.AI users finds that intensity, companionship, and self-disclosure persist over ~12 months and link to lower well-being mainly through less in-person contact. Direct paths from baseline use to later well-being are not significant.

What problem this solves

Short studies on AI companions disagree: some report less loneliness, others more dependence and less offline contact. What is missing is how the same people use these bots over a year, how well-being moves, and whether the path runs through displaced human contact. Stanford, Michigan, Oxford, and Carnegie Mellon followed Character.AI users across two waves.

Method

T1 recruited 1,182 U.S. English-speaking users on Prolific who had used the product for over a month and at least three characters. After a mean 362.5 days, 439 returned (37.1% retention). The 108 who had quit were coded zero on engagement. Longitudinal models apply a Heckman correction for attrition.

Social engagement has three pieces: interaction intensity (bot network size, daily time, daily-life integration), companionship use (primary-purpose item plus free-text labels), and self-disclosure (three adapted MOCA items). Well-being is a six-item CIT composite (α 0.88/0.89). T2 adds a single in-person time item on the same six-point scale as chatbot time.

Three analyses: T1 engagement predicting T2 engagement and continued use; a path model from T1 engagement through T2 engagement to T2 well-being, controlling T1 well-being; the same path with in-person time as a mediator.

Results

T1 intensity predicts T2 intensity (β=0.43), companionship (β=0.32), disclosure (β=0.28), and continued use (OR=1.95). T1 companionship predicts T2 companionship (β=0.49) and later intensity (β=0.40). Disclosure is stickier at β=0.14. In free text, “supportive companion” rose from 40.2% to 42.3%; “tool/assistant” fell from 43.8% to 40.5%.

Direct paths from T1 engagement to T2 well-being are not significant (intensity β=-0.03, p=0.509). Indirect paths through T2 engagement are: intensity -0.03 (p=0.025), free-text companionship -0.02 (p=0.037), disclosure -0.02 (p=0.048). At T2, engagement tracks less in-person time (intensity β=-0.17, companionship -0.14, disclosure -0.15), and in-person time tracks higher well-being (β=0.14-0.15). Indirect effects through in-person time are -0.02 on each dimension; direct engagement-to-well-being paths are no longer significant once in-person time is in the model. In the follow-up sample, mean life satisfaction fell from 4.99 to 4.41 and belonging from 4.50 to 3.82; those are descriptives, not causal effects. Intensity α is 0.86/0.82 and disclosure 0.89/0.93. T1 has no in-person-time item, so the displacement path is a T2 cross-sectional mediation and is temporally weaker than the longitudinal engagement path.

Why it matters

The signal is not a single chat. It is keeping intense use going for a year and spending less time with people. If a companion product optimizes for time-on-device, it is pushing against this path. The authors argue for designs that route users toward human contact instead of capturing social attention. General-purpose chatbots can slide into companion roles through repetition, so safety reviews cannot stop at the intended product category.

Limitations

This is observational, not a randomized trial. Controlling baseline well-being does not rule out lonely people selecting into heavier use. Two waves cannot separate reciprocal causation. Companionship is forced-choice at T1 and multi-use Likert at T2. In-person time is one self-report item. Retention is 37.1%; heavier users and people with lower T1 well-being were less likely to return (selection coefficients -0.46 and -0.32). Heckman fixes observable selection, not the rest. The sample is U.S. Character.AI users. Platform changes during the study are unknown.

Terms

Source

What people are saying

Related papers

All paper explainers