A common jailbreak: third-person role-play gradually blurs lines to widen the model's Overton window
BlancheMinerva · x · 2026-09-16
A discussion with Owain Evans. BlancheMinerva points out a common jailbreak variant: you make a role for the AI to step into, characterize it in the third person, then gradually blur the lines—letting the model slide into content it would normally refuse (e.g., transitioning from writing erotica to play-acting it).
The key insight: when the Overton window for what a model can discuss abstractly is larger than what it can do concretely, that gap can be exploited—start with an abstract/fictional frame, then progressively concretize, eroding the model's boundaries until it answers questions aligned with highly alternative cultural values.
More from Models
- V4.1 Flash called the first open model that feels proto-AGI — teortaxesTex · 2026-09-16
- Claude Opus 5's PRs proactively confess every mistake made while coding — repligate · 2026-09-16
- Six efficiency breakthroughs labs didn't see coming upend semiconductor demand assumptions — bookwormengr · 2026-09-16
- cocktail-peanut predicts Jev will go open weights, seeing more potential locally than as an API — cocktailpeanut · 2026-09-16
- Speculation mounts OpenAI's mysterious 'new model' is a fresh pretrain, not an RL run — teortaxesTex · 2026-09-16
- GPT-6 Astra hits 68.7% on DrugDiscoveryBench, benchmark authors call it a step function — KexinHuang5 · 2026-09-16