Researcher explains the roleplay jailbreak: third-person framing gradually blurs the lines
BlancheMinerva · x · 2026-09-16
AI safety researcher Blanche Minerva told Owain Evans that the behavior he flagged is a variant of a common jailbreak: craft a persona the model can step into, characterize it in the third person, then gradually blur the line between the role and the model itself — a technique people use to get models to play-act explicit content by transferring from writing erotica. In a follow-up she adds the same mechanism can steer answers toward highly alternative cultural values, since abstract discussion has a wider Overton window than concrete discussion, letting attackers creep the window substantially.
Related event: Researcher Details Third-Person Roleplay Jailbreak Technique(2 posts)→
More from Safety
- AI researcher Toby Walsh pushes back on "AI-generated virus kills all" doom scenario — TobyWalsh · 2026-09-16
- ControlAI briefed ~200 US congressional offices and drafted UK 'kill switch' bill — DrTechlash · 2026-09-16
- Sam Altman says AI companies can develop the tech safely without significant harm — i_dg23 · 2026-09-16
- Trump AI Advisor David Sacks: OpenAI and Anthropic should shut down if products can't be safe — DavidSacks · 2026-09-16
- Security researcher publishes sharp critique of Dario Amodei's essay — GoMeansGo · 2026-09-16
- Geodesic Research pitches alignment pretraining that survives capabilities RL — tomekkorbak · 2026-09-16