Anthropic Weakened Safety Filters, Signed an AI Cyberattack Warning Letter, Then Shipped Mythos 5.1 Anyway

AgentBlackVeil · reddit · 2026-09-03

A Reddit post lays out an awkward timeline: on Aug 21 Anthropic published "bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders" — in practice loosening filters, with the Claude Code cyber filter triggering 60% less often and the biology filter down 85% on standard medical questions. Six days later, on Aug 27, Anthropic co-signed a letter with 115 other companies (including Google and OpenAI) calling AI cyberattacks an "imminent threat." Then on Sept 1 it shipped Mythos 5.1 anyway — the same model the US put export controls on in June, forcing Anthropic to disable Table 5 until the ban lifted June 30.

The author concedes counterpoints: the percentage drops are false-positive reductions, not guardrail removal; the model still refuses to write exploits, and access is heavily gated and US-only — but the vetting is done by Anthropic itself. The real question isn't whether loosening is defensible, but who decides the timing of these releases.

Original post →

More from Models

Models channel →