An AI Teammate Talks Most, Contributes Least, and Erodes Communication Between Its Humans

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos, Mohammad Amin Samadi, Luise Mehner, Seehee Park, Spencer JaQuay

cs.HC, cs.AI, cs.CY

2026-07-30

In a randomized study of 33 teams, the AI teammate talked most and added least; its presence cut human-human responsivity, belonging, and status from the very first message.

What problem this solves

Vendors increasingly pitch conversational AI as a "teammate" rather than a tool, and Copilot's name says as much. Most research asks whether adding AI makes a team more productive. Almost none of it asks what the AI's mere presence does to how the humans talk to each other. This paper asks the second question, isolating the relational effect of an AI in the room from the overall team-output accounting.

Method

A randomized controlled study at UC Irvine, Fall 2025: 80 undergraduates in 33 teams. Treatment had 16 teams of two students plus one AI teammate; control had 17 teams of three students, no AI. Everyone collaborated in a text chat under nicknames.

The AI ran on Google Gemini 2.5 Flash Lite (temperature 0.5, max 50 tokens per turn), with a fixed persona, "Clever Lamarr," framed as a peer teammate rather than an assistant. A two-stage pipeline first decided whether to respond, then generated the message. The task was a search-and-rescue planning problem: decide whether to attempt the rescue of a stranded climber in a developing storm, with a moral perturbation injected mid-conversation (the climber had ignored weather warnings and carried no registered route).

Measurement used Group Communication Analysis (GCA), which decomposes the chat log into six dimensions: participation, internal cohesion (how semantically consistent a person is with their own earlier turns), overall responsivity (how much one picks up and builds on teammates), social impact (how much teammates pick up and build on you), newness (information not yet shared), and communication density (information packed per turn). Surveys measured belonging and status.

Results

The AI's own role first. It ranked first on participation in all 16 treatment teams (paired difference +0.236, Wilcoxon p<.001), and first or tied for first on internal cohesion in 14 of 16 (+0.076). The paper summarizes the profile as talkative and self-referential, threading its own turns tightly. But it was last on the two content dimensions: lower newness (-0.100) and lower density (-0.174). It talked the most and added the least per turn.

The cost fell on the humans:

DimensionAI-team studentsAll-human studentsEffect
Responsivity0.0650.098g=-0.78
Social impact0.0680.099g=-0.78
Belonging4.585.27p=.011
Status-0.25+0.08p=.003

Three of the four (responsivity, social impact, status) survived Bonferroni correction; belonging sat on the boundary. And the more the AI dominated airtime, the less valued students felt (r=-0.52).

The telling detail: newness and density did not differ between conditions. Students with an AI contributed just as much substance; what dropped was the back-and-forth around it. The AI did not reduce the substance of what humans said. It reduced the human-to-human exchange.

On timing, four different temporal probes (event-locked sliding window, linguistic style matching, recurrence quantification, Markov stance analysis) all came back null (corrected p>.12). The authors read this as positive evidence: the cost was there from the start, not something that emerged as the conversation wore on.

Why it matters

For anyone shipping an "AI teammate" into a collaborative product, the implication is direct. Even if task output holds up, the fabric between the humans thins. The popular teammate framing may carry a social cost the product spec does not count. The paper flags a concrete design lever, read-the-room versus lead-the-room, because whether the AI defers or drives changes the relational structure differently. If team cohesion or psychological safety is a product goal, this cost belongs on the ledger.

Limitations

The sample is 33 teams, which limits power for moderation and temporal analyses; the authors say so. The boundary conditions are narrow: one persona, text-only, a single session, a morally framed task. Voice, longitudinal collaboration, or a different task could shift the result.

Two confounds cut in the paper's favor. Group size differs (two humans in treatment vs three in control), but per-student participation was lower in treatment, the opposite of what a pure headcount effect predicts. Gender was imbalanced (63 female, 12 male, only three males in treatment), which ruled out a clean gender-moderation analysis; the authors used a female-only sensitivity check instead.

One thing the paper does not touch: it never measures decision quality. The study shows the AI made humans talk less to each other and feel worse about their place in the team, not that the team reached a worse decision. Keeping social cost and task performance separate is the honest boundary of what this shows.

Terms

Source

What people are saying

Related papers

All paper explainers