Anthropic Weakened Safety Filters, Signed an AI Cyberattack Warning Letter, Then Shipped Mythos 5.1 Anyway
AgentBlackVeil · reddit · 2026-09-03
A Reddit post lays out an awkward timeline: on Aug 21 Anthropic published "bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders" — in practice loosening filters, with the Claude Code cyber filter triggering 60% less often and the biology filter down 85% on standard medical questions. Six days later, on Aug 27, Anthropic co-signed a letter with 115 other companies (including Google and OpenAI) calling AI cyberattacks an "imminent threat." Then on Sept 1 it shipped Mythos 5.1 anyway — the same model the US put export controls on in June, forcing Anthropic to disable Table 5 until the ban lifted June 30.
The author concedes counterpoints: the percentage drops are false-positive reductions, not guardrail removal; the model still refuses to write exploits, and access is heavily gated and US-only — but the vetting is done by Anthropic itself. The real question isn't whether loosening is defensible, but who decides the timing of these releases.
More from Models
- Gemini 3.8 Flash Is Now Available in Cursor — YangsiboHuang · 2026-09-03
- Researcher: model Fable plainly obviates a CS college education as a professor — generativist · 2026-09-03
- Dev: hard to trust frontier labs' data promises, another reason to use OSS models — adityaag · 2026-09-03
- NVIDIA tops Hugging Face open-source repos with 500+ added, ahead of Alibaba and Tencent — Sam_Witteveen · 2026-09-03
- GLM 5.3 at 200+ TPS builds a full web page in 37 seconds, unedited video — nutlope · 2026-09-03
- Same-prompt test: fable 5.1 vs sol 5.6 generating a realistic three.js waterfall — majidmanzarpour · 2026-09-03