Paper: Do Models Believe What They Say When Role-playing?

kastnerkyle · x · 2026-08-27

A new paper investigates whether LLMs internally adopt the beliefs of the personas they roleplay. Researchers had models portray historical characters with views contradicting modern consensus and used probes to check if the models' internal belief states matched their (historically appropriate but factually wrong) statements.

Results show a spectrum of belief internalization across methods (Prompting, ICL, SFT, Open Character Training, Emergent Misalignment). Prompting, ICL, and SFT change outputs with little representational shift. Emergent Misalignment causes a broad, significant shift in the model's truth representation, while Open Character Training causes a smaller, clearer shift in larger models.

Original post →

More from Research

Research channel →