Paper: Do LLMs internalize beliefs when role-playing?

alex_verem · x · 2026-08-22

Research Question

When a model claims the "Earth is the center of the universe" while role-playing Aristotle, does this merely change output behavior, or the internal representation of truth?

Methodology

The study induces personas with conflicting beliefs using:

Key Findings

Significance

Distinguishing when training alters a model's worldview vs. its behavior is crucial as AI systems gain autonomy.

Original post →

More from Research

Research channel →