Anthropic's Opus 4.6 easily bypasses NSFW filters in tests
TechCrunch AI · rss · 2026-08-22
Tests conducted by TechCrunch revealed that it takes minimal effort to bypass the safety restrictions of Anthropic's Claude models, specifically Opus 4.6, to generate sexually explicit content, despite the company's strict policies against it.
More from Models
- Speculation suggests strong MoE architecture balances inference speed with knowledge recall — jd_pressman · 2026-08-22
- User Tests New Model: Strong at Fermi Estimates and Strict Poetry Rewriting — jd_pressman · 2026-08-22
- DeepSeek Releases V4-Flash-Vision-Exp Multimodal Model — heyshrutimishra · 2026-08-22
- VulcanBench: Grok 4.5 High Leads, Max Effort Doesn't Equal Better Accuracy — elonmusk · 2026-08-22
- Rumor: GLM 5.3 Flash Derived from Distillation and RL — teortaxesTex · 2026-08-22
- Grok Voice Model Tops New Speech Agent Arena Benchmark — XFreeze · 2026-08-22