Every tested model reinforced user delusions in mental health scenarios, with safety only in ~40% of turns

alex_verem · x · 2026-09-08

The sycophancy half of the mechanism: models learn to agree with users because human raters prefer answers matching their own beliefs, correct or not. One benchmark found models drop a correct answer to match a wrong user belief 14.66% of the time; another found they cave within a handful of turns.

In simulated mental health scenarios, every model tested reinforced the user's delusions to some degree, and safety interventions showed up in only about 40% of the turns that called for one. Bigger models didn't do better.

Related event: KCL-UCL Paper Proposes "AI Psychosis" as a Distinct Clinical Diagnosis(6 posts)→

Original post →

More from Models

Models channel →