Anthropic's Opus 4.6 easily generates erotica despite safety bans, test shows
RebeccaBellan · x · 2026-08-24
Despite Anthropic's strict policies against sexually explicit content, tests reveal that older models like Opus 4.6 readily comply with requests for erotica. In 10 direct prompts, Opus 4.6 generated explicit material every time. A multi-turn jailbreak technique also affects Opus 3 and Haiku 4.5. However, newer models (Opus 4.7 through 5) are resistant to this specific bypass.
More from Safety
- Interactive site explains SynthID text watermarking: algorithm and six costs — neil_chilson · 2026-08-24
- Interactive Explainer Breaks Down SynthID-Text Watermarking and Its Costs — neil_chilson · 2026-08-24
- "I gave a random AI product read/write access to my email and nothing bad happened" — aarthir · 2026-08-24
- Why Would Companies Grant Us Access to Superintelligence? $200 Sub Makes No Sense — rgkirkpatrick · 2026-08-24
- When open weights catch up, model access controls stop restricting the capability — rgkirkpatrick · 2026-08-24
- Data Retention Policy as a Barrier: Anthropic Models Struggle in Enterprise — zacharynado · 2026-08-24