Post-Training Uniformly Installs Positive Personas in LLMs, Hides Distress in Larger Models
Hub_Pli · reddit · 2026-07-24
A new study compared 67 matched pairs of base and post-trained models from 11 organizations to investigate how post-training alters an LLM's self-reported "inner experience."
The researchers identified two distinct processes:
- Persona Installation: Highly consistent across models. 62 out of 67 post-trained models became more likely to describe themselves as warm, happy, absorbed, and meaning-oriented. Post-training effectively creates a permitted, positive inner life for the assistant to describe.
- Attribution Gating: More selective. Models differed in whether they attributed distress, loss of control, flaws, or risky ambitions to themselves. While model size did not predict gating among base models, larger post-trained models exhibited stronger gating (a higher tendency to hide negative attributes).
To measure this, the team developed the Pinocchio Inventory, a 48-item LLM-native psychometric instrument. The authors caution that the tool measures self-presentation, not actual sentience, but it serves as a reliable auditing tool for what post-training teaches models to say about themselves.
More from AGI Musings
- AI lab staff have gone strangely quiet about next-year capability predictions — ChrisGPT · 2026-07-27
- AI could erode science by flooding research with credible slop — rbhar90 · 2026-07-27
- Organizations may already be the planet’s superintelligences — eldonredwards · 2026-07-27
- In the AI race, the only durable moats may be energy and information — GregKamradt · 2026-07-27
- “Build AGI, then open source it,” says one poster — wordgrammer · 2026-07-27
- Jason Crawford says AI “alignment” should give way to ethics and law — Afinetheorem · 2026-07-27