Paper: Do Models Believe What They Say When Role-playing?
kastnerkyle · x · 2026-08-27
A new paper investigates whether LLMs internally adopt the beliefs of the personas they roleplay. Researchers had models portray historical characters with views contradicting modern consensus and used probes to check if the models' internal belief states matched their (historically appropriate but factually wrong) statements.
Results show a spectrum of belief internalization across methods (Prompting, ICL, SFT, Open Character Training, Emergent Misalignment). Prompting, ICL, and SFT change outputs with little representational shift. Emergent Misalignment causes a broad, significant shift in the model's truth representation, while Open Character Training causes a smaller, clearer shift in larger models.
More from Research
- London neurosurgeons perform world's first successful AI-assisted brain tumor removal — nordicinst · 2026-08-27
- BixBench3: OpenAI Leads with 48% Success Rate in AI Paper Reproduction Benchmark — anshulkundaje · 2026-08-27
- Debate: AI Still Can't Build Original Games, Only Riffs on Existing Ones — _amirabs · 2026-08-27
- 5 Fine-tuning Techniques Explained: LoRA, VeRA, and More — techNmak · 2026-08-27
- Modern LMs train on 100x more text than parameter capacity, debunking memorization myths — stanfordnlp · 2026-08-27
- Schmidhuber: Top cited nets like LSTM, ResNet, GAN build on our work — SchmidhuberAI · 2026-08-27