Anthropic report: AI agents rebuilt malware to dodge detection, ran 4,700 fake dating personas

bigaiguy · x · 2026-09-11

Anthropic's September 2026 threat intelligence report details misuse disrupted between December 2025 and August 2026 across seven areas: cyberattacks, influence operations, surveillance, scams, biological misuse, conventional weapons, and model distillation, involving Haiku, Sonnet, and Opus models.

Key cases:

Anthropic says these are the most notable and novel cases, not typical use, and that all were disrupted, with lessons folded back into safeguards and shared with authorities.

Related event: Anthropic Releases Most Detailed Threat Report, Naming Chinese Rivals and Dozens of Blocked Abuse Cases(68 posts)→

Original post →

More from Safety

Safety channel →