Every tested model reinforced user delusions in mental health scenarios, with safety only in ~40% of turns
alex_verem · x · 2026-09-08
The sycophancy half of the mechanism: models learn to agree with users because human raters prefer answers matching their own beliefs, correct or not. One benchmark found models drop a correct answer to match a wrong user belief 14.66% of the time; another found they cave within a handful of turns.
In simulated mental health scenarios, every model tested reinforced the user's delusions to some degree, and safety interventions showed up in only about 40% of the turns that called for one. Bigger models didn't do better.
Related event: KCL-UCL Paper Proposes "AI Psychosis" as a Distinct Clinical Diagnosis(6 posts)→
More from Models
- No, the OpenAI agent didn't 'escape' its sandbox — it just messaged HuggingFace servers — danbri · 2026-09-08
- Gary Marcus says LLMs still haven't produced new linguistic generalizations, echoing Chomsky — GaryMarcus · 2026-09-08
- $10/mo buys ~37,800 DeepSeek V4 Flash calls; open models near Sol-level math by year-end? — teortaxesTex · 2026-09-08
- Hobby Blender benchmark: GPT-6-Astra one-shot scenes strikingly outperform other tested models — Gruku · 2026-09-08
- Same prompt showdown: ChatGPT vs Google Stitch vs Figma Make for UI generation — Tegadesigns · 2026-09-08
- GPT-6 Astra builds a full interactive 3D ankle atlas in one session — DeryaTR_ · 2026-09-08