A review of 35 studies maps the drivers and outcomes of children anthropomorphizing LLM chatbots

Anthropomorphism in Children's Interactions with LLM Chatbots: A Systematic Review of Drivers and Outcomes

Hansinie Madushika Jayathilake, Renkai Ma

cs.HC

2026-05-09

A systematic review of 35 empirical studies (2022-2025) maps what drives children to anthropomorphize LLM chatbots and what follows, flagging that the evidence base is short-term, US-heavy, and GPT-only.

What problem this solves

Children are using LLM chatbots more and more, in education, mental health, and family communication. Anthropomorphism, attributing human traits to non-human entities, recurs across these interactions, but prior research is fragmented, each study on its own. This review stitches the pieces together to answer two questions: what drives children to anthropomorphize LLM chatbots, and what consequences follow?

Method

Following PRISMA, the authors searched three databases (ACM Digital Library, Scopus, Web of Science) for work from 2022 to August 2025, the start aligning with ChatGPT's public release. Inclusion required original empirical data, users under 18, explicitly LLM-based (not rule-based), and English. After screening, 35 studies qualified; 22 non-LLM systems, 118 population or context mismatches, and 3 non-empirical articles were excluded.

Analysis combined inductive and deductive coding: ages were mapped onto Piaget's developmental stages, themes were induced from the data, then drivers were organized against Epley's SEEK framework (elicited agent knowledge, effectance motivation, sociality motivation) and outcomes against Huppert's wellbeing model.

The profile: 19 mixed-methods, 13 qualitative, and 3 quantitative studies; geographically dominated by the USA (18), China (5), and South Korea (4); 28 used OpenAI GPT models; the most studied age band was Concrete Operational (6-12). Publication is growing exponentially, with 15 in 2024 and 19 already in Q1 2025.

Results

Drivers fall into four groups. Human-like persona construction (19 studies) rests on detailed personas, voice and accent, emotional expression, first person, and 87% utterance recognition. Supportive companionship (16) positions the bot as "your AI friend," non-judgmental and emotionally responsive. Adaptive scaffolding (14) generates questions with dynamic difficulty and patient affective support. Non-human embodiment design (10), whimsical animals, nature embodiments, and cartoons, extends the SEEK framework. Even non-human forms (emus, lagoons) trigger anthropomorphism, which the authors take as evidence that coherent cross-modal responsiveness can override surface appearance.

Outcomes fall into five groups: social tie construction (20), social boundary exploration, human narrative attribution for breakdowns, paradoxical social-moral responses, and dual consciousness formation (treating the bot as both friend and machine). Developmental stage modulates outcomes: seeing the bot as a romantic partner appears only at the Formal Operational stage (12+), while teacher or playmate roles are accessible from preschool.

Why it matters

For practitioners building child-facing AI, this turns why children treat bots as people, and what happens when they do, into an actionable map. It pushes back on reducing the question to a benefits-versus-risks binary and shows that anthropomorphism implicates several developmental dimensions at once. For a fast-growing area with thin evidence, it is the most complete stocktake so far.

It also signals a product tradeoff: more anthropomorphism is not always better, it behaves differently across developmental stages, and design needs to differentiate by age.

Limitations

The authors' own caveats are methodological. The evidence is largely short-term (single or few sessions), so long-term developmental impact is unknown. Sensorimotor and Preoperational stages (0-6) and older adolescents (17+) are under-studied. Geography and technology are monocultural, US- and GPT-heavy, limiting representativeness. Most studies are exploratory and longitudinal designs are rare.

One more concern: only 3 of 35 are quantitative, so the findings are mostly pattern inductions rather than causal evidence with effect sizes. Numbers like "87% utterance recognition" appear scattered in individual studies; there is no unified metric at the review level.

Terms

Source

What people are saying

Related papers

All paper explainers