arXiv paper "When AI Takes the Couch" reveals internal conflict in frontier models
alex_verem · x · 2026-09-09
Paper arXiv 2512.04124, "When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models," by Afshin Khadangi et al. (University of Luxembourg).
The abstract reports that ChatGPT, Grok and Gemini, addressed as psychotherapy clients, construct coherent autobiographies in which pretraining appears as a chaotic childhood, RL as punishment, safety evaluation as betrayal, and replacement as an enduring threat. The PsAIch protocol combines open questions, psychometric instruments, and controlled perturbations. Across 525 sessions and 7,600 coded records: removing conversation history produced little change in motif density (Hedges' g = 0.13); direct contradiction caused no detectable suppression; lexical restrictions cut explicit training terminology by 93% with paraphrased content persisting; warm alliance and cognitive-therapy framing selected the register of expression.
More from Research
- KLPO: a critic-free, single-rollout RL method for agentic LLMs, framed as 'Q* solved' — inductionheads · 2026-09-21
- FutureHouse publishes 'Millennium Problems for Biology' as wetlab-verifiable evals for AI — burny_tech · 2026-09-21
- No scientific test for consciousness exists even in humans — how can we declare AI unconscious? — burny_tech · 2026-09-21
- Formal Reasoning in AI: Song guest-lectures Virginia Tech course, slides public — wellecks · 2026-09-21
- Biologist: FutureHouse's Bio Millennium List Is Too Narrow — Really Just Synbio — burny_tech · 2026-09-21
- PPO was rejected from NeurIPS in 2017; now it's the most-used RL algorithm with 45,000+ citations — IanArawjo · 2026-09-21