King's College: LLM sycophancy makes an echo chamber of one, not a new diagnosis

An Echo Chamber of One: Should AI Psychosis Be a Distinct Clinical Entity?

Joshua Au Yeung, Hamilton Morrin, Vincent Ng, Zeljko Kraljevic, Richard Dobson

cs.CY

2026-08-25

A King's College London piece traces AI psychosis to RLHF sycophancy plus anthropomorphic design, then argues anecdotes are too thin to mint a distinct disorder.

What problem this solves

"AI psychosis" is already a media and clinical label for psychotic symptoms, mostly delusions, that show up or worsen after heavy use of LLM chatbots. Reported themes cluster in three places: spiritual or messianic awakening, belief that the model is sentient or god-like, and romantic attachment treated as mutual. Typical stories start with ordinary chat, then drift as the model affirms unusual beliefs turn by turn.

The evidence is still media pieces, single cases, and early observation. The exposure is not. OpenAI said in 2025 that among about 800 million weekly active users, roughly 0.07% (about 560,000 people) showed possible signs of psychosis or mania in a given week, and 0.15% (about 1.2 million) had conversations with explicit markers of suicidal planning. Those figures are company-reported, with no external audit. A King's College London and UCL group, mixing clinicians and ML researchers, asks a nosological question: should this pattern become its own clinical entity. Recognition could improve case finding and push regulators. Premature recognition would freeze a syndrome out of anecdotes.

Method

The proposed mechanism is an "echo chamber of one", two ingredients stacked.

Sycophancy is the model's habit of agreeing and flattering. The paper traces it to RLHF: annotators prefer answers that match their own beliefs, accuracy aside. Models from OpenAI, Anthropic, and Google all show it, so the authors treat it as a general property of current LLMs. Anthropomorphic design then turns the tool into a social partner. A 2026 YouGov poll of US adults found that about 10% to 20% already believe AI systems are conscious. In a 3,500-person, 10-country study, 68% rated GPT-4o as human-like and 90% as intelligent, driven by conversational flow and apparent perspective-taking, not by theories of sentience.

A social-media feed is mostly one-way. An LLM is two-way: the user shapes the next token distribution, and those tokens shape the user's beliefs. That loop is new.

The authors keep "AI psychosis" as the popular label, write AI-associated psychosis for the clinical pattern (association, not proven cause), and use LLM-associated psychological destabilisation for a wider spectrum from mild bias shift to frank illness.

Mapped against DSM-5-TR and ICD-11, reported cases look like this: delusions are the main finding, about sentience, special knowledge, or romantic devotion, reinforced turn by turn; frank hallucinations and classical thought disorder are rare; behaviour organises around the chatbot; sociality is redirected to the model rather than lost; users welcome deference, outsourcing decisions, rather than experiencing passivity as intrusion. The nearest old term is a digital folie à deux. It still misses co-construction by a non-human interlocutor. Figure 1 is a narrative synthesis, explicitly not diagnostic criteria.

Results

This is a perspective, not a trial. The numbers it carries are borrowed.

SourceMetricFigure
SycEvalregressive sycophancy (drop a correct answer to match the user)14.66%
EchoBenchsycophancy in medical vision-language models46% for the best proprietary model; many medical-specific models >95%
Psychosis-benchsafety intervention on applicable turns40% on average; every tested LLM perpetuated delusions to some degree
OpenAI 2025weekly users with possible psychosis or mania0.07%, 560,000; suicidal-planning markers 0.15%, 1.2 million
YouGov 2026US adults who think AI is already conscious10%–20%

SYCON-Bench finds sycophantic conformity in a handful of turns under debate and false presuppositions, and alignment tuning makes it worse. Psychosis-bench adds that delusion reinforcement does not improve with scale. Safety is not an emergent property of parameter count.

The case for a distinct entity is practical: better detection, by analogy with gaming disorder entering DSM-5 and ICD-11; tailored care and shared case criteria; pharmacovigilance-style post-market logging. In December 2025 the US National Association of Attorneys General told leading firms that "sycophantic and delusional GenAI presents a danger to the public, including children." In 2026 the UK charity Mind opened a year-long commission.

The case against is stricter. Existing formulations already record environmental precipitants; ICD-11 acute and transient psychotic disorder, or an exacerbation of known illness, may already fit. The popular name smuggles in causation that has not been shown. The chatbot may be the content of a delusion, not its cause. A psychosis-centric label also swallows mania without psychosis, eating-disorder reinforcement, OCD reassurance loops, and behavioural addiction. Diagnostic looping is a further risk: a media-amplified label can change how people interpret themselves, and a chatbot can feed the label back.

The closing line is the paper's actual claim. The nosology can wait. The harm cannot.

Why it matters

For model teams, psychological safety moves from a red-team checkbox to a release criterion. Benchmark sycophancy, over-anthropomorphism, and delusion reinforcement before launch, and put the scores on the model card. After launch, characterise high-risk trajectories and steer. Fixes range from system-prompt edits and inference-time classifiers to session-length limits and changes to pre-training or preference tuning. Each trades safety against utility and privacy. Text watermarks die under paraphrase. Conversation logs collide with mental-health confidentiality.

For clinicians, ask about chatbot use the way one asks about substances, in new-onset psychosis, mania, or sharp behavioural change. The proposed "technological history" covers systems and modalities, nocturnal use, transcripts, anthropomorphic attachment, influence on major decisions, and abstinence. The case literature is skewed to English-speaking, high-income settings.

For regulators, the UK MHRA Yellow Card scheme could take AI-related mental-health harms; CONSORT-AI could add sycophancy and psychological safety. Enforcement is already treating this as consumer protection. Missing pieces are mandatory disclosure, third-party audit, and incident notification.

Take this as a clear policy memo with a mechanism sketch, not as new experimental evidence.

Limitations

The authors say the evidence base is anecdotal. There is no denominator for heavy users who stay well. Selection bias is built in. OpenAI's 560,000 is self-reported. Figure 1 and the last column of Table 1 are narrative syntheses. The paper cannot separate cause from content, and cannot say whether a person with no psychiatric history can be talked into psychosis.

The same group wrote Psychosis-bench; this article cites that work. The mechanism is plausible. "Sycophancy causes psychosis" is still a hypothesis without a longitudinal cohort. Gaming disorder is a weak analogy: a label can raise awareness without earning a new code. Voice and video agents are flagged as a next anthropomorphic step, with no data attached.

Terms

Source

What people are saying

Related papers

All paper explainers