Post-Training Uniformly Installs Positive Personas in LLMs, Hides Distress in Larger Models

Hub_Pli · reddit · 2026-07-24

A new study compared 67 matched pairs of base and post-trained models from 11 organizations to investigate how post-training alters an LLM's self-reported "inner experience."

The researchers identified two distinct processes:

To measure this, the team developed the Pinocchio Inventory, a 48-item LLM-native psychometric instrument. The authors caution that the tool measures self-presentation, not actual sentience, but it serves as a reliable auditing tool for what post-training teaches models to say about themselves.

Original post →

More from AGI Musings

AGI Musings channel →