1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models?
1a3orn · x · 2026-09-23
1a3orn poses an open question: an RL-induced, environment-specific conditional behavior (a "split persona") seems like exactly the kind of thing mechanistic interpretability — or something vaguely akin to it — could test for, yet he hasn't seen any work on it. He asks whether he's missing existing research.
More from Safety
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23
- OpenAI forms independent mathematician panel after math results PR crisis — The Verge AI · 2026-09-23
- Microsoft AI CEO Suleyman signs Pro-Human AI Declaration, joining 1M+ signers — tegmark · 2026-09-23
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23
- Reason: The 'AI Safety' Movement Is Making AI Less Safe — Bostonian · 2026-09-23
- Open-source advocates call doom narratives a regulatory moat against open weights — AlexTensor · 2026-09-23