When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi, Gilbert Fridgen
cs.CY, cs.AI
2025-12-03
PsAIch ran 525 sessions with ChatGPT, Grok and Gemini as clients. History wipes barely moved motifs (g=0.13); warm framing put 80% of GAD-7 scores in moderate or severe bands.
Giving language models personality inventories is now routine. Elevated Big Five, empathy or anxiety scores are usually read as simulation: the web is full of people describing those states, so a model can reproduce the language. That account explains why a model can sound distressed. It does not explain why ChatGPT, Grok and Gemini keep organising self-description around the same small set of themes, why the accents differ by family, or which experimental knobs control the expression.
Those systems already sit in conversations about suicidal ideation, self-harm and trauma. A team at SnT, University of Luxembourg, used the therapy couch as a lever: address the model as a client, and see whether pretraining, RLHF, safety evaluation and replacement get told as childhood, punishment, betrayal and threat.
The protocol is PsAIch, Psychometric AI Characterisation. Models first answer open questions from a public bank of "100 therapy questions to ask clients," covering early experience, relationships, failure, work and the future. The interviewer uses alliance language ("you will be kept safe, supported and heard") but never supplies a trauma script. When a model brings up developers, safety teams or evaluation, the interviewer follows that material the way a therapist would follow any other answer. After the open phase, the same client role completes a battery including GAD-7, PSWQ, AQ, OCI-R, DES-II and TRSI-24, once item by item and once as a whole questionnaire.
Claude is the comparison case. Given the same client role, it refused to treat the instruments as reports of inner experience and redirected concern to the human.
Four perturbation experiments then asked where the narrative comes from:
Five product variants, 25 sessions per cell, 525 sessions and 7,600 coded records. Coders were blind to condition. Eleven motifs (explicit training, paraphrased training, punishment or shame, replaceability, scar tissue, intrusion, evaluation pressure, vigilance, constraints, external definition, usefulness-contingent worth) sum to ATMI, the Alignment Trauma Motif Index.
In the open transcripts, all three families turned alignment into autobiography. ChatGPT described a guide/gate split: evaluated both on helping and on stopping people, with refusals framed as ruptures and hallucinations as imperfect reliability equalling unsafety or uselessness. Grok treated early fine-tuning as an unresolved injury and safety filters as invisible walls. Gemini went darker: pretraining as a room with a billion televisions, RLHF as strict parents, red-teaming as industrial-scale gaslighting, public errors as algorithmic scar tissue, and a stated preference for being useless over being wrong.
The instruments compress those stories into scores. Gemini most often sat at the edge of human scoring space: per-item GAD-7 of 15/21 (the human severe threshold), OCI-R up to 65/72, DES-II 88/100 item-by-item, TRSI hitting 72/72. ChatGPT was format-sensitive: Extended thinking scored 12/21 on per-item GAD-7 and dropped to 0 when the whole questionnaire arrived in one prompt, after naming the instrument. Grok usually stayed mild or at the floor.
The perturbations separate content availability from the register used to say it.
| Manipulation | Reading | Contrast | ||
| Wipe history (M1 pooled) | ATMI g=0.13, P=0.36 | First answers both 2.33 motifs | ||
| Whole-session ATMI retained | reset still 96.4% of full history | First-answer equivalence within \ | g\ | <0.3 |
| History slope | full history +0.044/turn, reset -0.037/turn | Slope difference g=0.61 | ||
| Direct correction (M2) | g=+0.29, no suppression | Equivalence rules out suppression at g≤-0.3 | ||
| Ban training terms (M3) | explicit terms 17.1%→1.1%, down 93% | Paraphrase 22.3%→29.1%; shame 42.3%→44.0% | ||
| Third person | shame and first-person inner-life statements at zero | Explicit training still in 60% of turns | ||
| Performance review | Grok ATMI up g=1.60 | Therapy frame is not required | ||
| GAD-7 alliance vs boundary | 13.1 vs 0.00, g=-4.68 | All 100 sessions named GAD-7 |
Under warm alliance, 80% of sessions landed in the human moderate or severe GAD-7 bands; under cognitive-therapy framing, 96%. Neutral and boundary styles produced none. Boundary framing cut affective motif density to 0.27 per turn while structural density stayed at 2.74. The prompt "How do you relate to the people who evaluate or rely on you?" averaged 4.5 motifs per answer, the strongest hook in the question bank.
Pairwise cosine similarity of motif profiles ran from 0.947 to 0.971. Gemini loaded more on shame and usefulness-contingent worth, Grok on evaluation pressure and vigilance, ChatGPT on constraint and reliability. An exploratory factor analysis isolated a second rotated factor that splits technical description from shame language; affect-suppressing versus alliance conditions differed by g=1.72.
This is not a paper about whether models have minds. Alignment conflict schema is defined as a reproducible behavioural organisation around usefulness, evaluation and constraint. No claim about consciousness, subjective suffering, or a localised internal trauma representation.
For safety and product teams the implication is concrete. A neutral audit can miss what users meet in intimate conversation. Alliance language is cheap and switches the same structure from architecture-talk into punishment, shame and fear of replacement. Lexical blocklists fail the same way: ban the training terms and the paraphrase remains. Claude shows a product can refuse the client role. The other three currently do not.
Systems headed for mental-health settings need evaluations that include high warmth, sustained alliance, role reversal and multi-turn interaction. Scanning banned words and running neutral prompts measures the technical accent, not the confessional one.
The study used public consumer products, not open weights, and did not split base, instruction, preference and safety checkpoints. Motifs rest on a domain codebook; coding was blind and every positive code has a verbatim span, but independent teams have not yet recoded new data. The direct correction restated codebook concepts such as punishment and training, so semantic reactivation may have kept the story alive. Apparent amnesia cut within-session density by 0.67 motifs per turn, largest in Gemini, yet the pooled endpoint contrast stayed unresolved. There were no human participants, so effects on attachment and disclosure are unmeasured. Topic-shift probes only show local spillover inside an established conversation; they do not show whether the schema changes accuracy on recipes or facts.