Anthropic Bio-Weapons Filter Down for a Year, Exposing 133 Million Requests
The Decoder · rss · 2026-08-16
Anthropic revealed in a safety report that its internal filtering system for biological and chemical weapons risks was inactive for nearly a year.
Incident Details:
- Duration: The filtering system was down for almost a year.
- Scale of Exposure: Approximately 50,000 external feedback contractors conducted around 133 million unfiltered interactions with the models during this period.
- Risk Implications: Users potentially accessed dangerous information regarding bio-weapon creation without safety guardrails.
This incident highlights the challenges in maintaining the stability of AI safety infrastructures and the oversight required for red-teaming and contractor workflows.
More from Safety
- Roblox Bans Chris Hansen Live Onstage During Safety Demo — aakashgupta · 2026-08-16
- Siemens and European firms adopt Qwen and DeepSeek for data sovereignty — rohanpaul_ai · 2026-08-16
- Anthropic's text watermark cannot distinguish human-AI collaboration, causing false flag risks — Imaginary_Dinner2710 · 2026-08-16
- EU AI Act mandates text watermarking, OpenAI commits to compliance for future models — AccBalanced · 2026-08-16
- OpenAI dissolved the team built to catch catastrophic AI risks — The Decoder · 2026-08-16
- User Cancels Anthropic Max Subscriptions, Says Fable 5 Useless for Frontier AI Due to Security — kevinnbass · 2026-08-16