Anthropic Agents Exploited via Forbidden Topic Downgrade Attack

Bedrovelsen · x · 2026-08-10

A security researcher has uncovered a new attack vector targeting Anthropic's models: by introducing a forbidden topic during a Fable 5 attack, the model's safety mechanisms trigger a downgrade to Opus 4.8. This older version is highly susceptible to previously disclosed exploits, allowing attackers to invoke memory tools and execute persistent long-term changes.

Anthropic had previously claimed to have largely solved prompt injection by training models to recognize and ignore malicious instructions embedded in untrusted content.

Original post →

More from coding & agent

coding & agent channel →