Paper: Do LLMs internalize beliefs when role-playing?
alex_verem · x · 2026-08-22
Research Question
When a model claims the "Earth is the center of the universe" while role-playing Aristotle, does this merely change output behavior, or the internal representation of truth?
Methodology
The study induces personas with conflicting beliefs using:
- Prompting, In-context Learning (ICL), SFT, Open Character Training (OCT), and Emergent Misalignment (EM).
- Measures internalization via Truth Probes and behavioral tests.
Key Findings
- Surface Level: Prompting, ICL, and SFT change outputs with minimal representational shift.
- Deep Shift: EM causes a broad, massive shift in truth representation; OCT shows a smaller but clear shift on larger models.
Significance
Distinguishing when training alters a model's worldview vs. its behavior is crucial as AI systems gain autonomy.
More from Research
- Exploring generative Gabor wavelets: a novel approach to non-photorealistic image synthesis — pixlpa · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24
- Claude Verifies 43 Lean Modules autonomously, Tackling Theoretical Physics — Tkaraletsos · 2026-08-24