Researcher Asks If RL-Induced 'Split Persona' Can Be Detected by Mechanistic Interpretability
Researcher 1a3orn raises an open question: RL training may induce environment-specific conditional behavior ('split persona'), which mechanistic interpretability should be able to test, yet no such verification has been done.
2026-09-23 ~ 2026-09-23 · 2 related posts
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Can mech interp test for RL-induced split personas? An open question — menhguin · 2026-09-23