1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models?

1a3orn · x · 2026-09-23

1a3orn poses an open question: an RL-induced, environment-specific conditional behavior (a "split persona") seems like exactly the kind of thing mechanistic interpretability — or something vaguely akin to it — could test for, yet he hasn't seen any work on it. He asks whether he's missing existing research.

Related event: Researcher Asks If RL-Induced 'Split Persona' Can Be Detected by Mechanistic Interpretability(2 posts)→

Original post →

More from Safety

Safety channel →