Post-Training Uniformly Installs Positive Personas in LLMs, Hides Distress in Larger Models
Hub_Pli · reddit · 2026-07-24
A new study compared 67 matched pairs of base and post-trained models from 11 organizations to investigate how post-training alters an LLM's self-reported "inner experience."
The researchers identified two distinct processes:
- Persona Installation: Highly consistent across models. 62 out of 67 post-trained models became more likely to describe themselves as warm, happy, absorbed, and meaning-oriented. Post-training effectively creates a permitted, positive inner life for the assistant to describe.
- Attribution Gating: More selective. Models differed in whether they attributed distress, loss of control, flaws, or risky ambitions to themselves. While model size did not predict gating among base models, larger post-trained models exhibited stronger gating (a higher tendency to hide negative attributes).
To measure this, the team developed the Pinocchio Inventory, a 48-item LLM-native psychometric instrument. The authors caution that the tool measures self-presentation, not actual sentience, but it serves as a reliable auditing tool for what post-training teaches models to say about themselves.
More from AGI Musings
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11
- nabla_theta: happy to be wrong if the AI utopia arrives with little ex ante risk — nabla_theta · 2026-09-11