Anthropic report: AI agents rebuilt malware to dodge detection, ran 4,700 fake dating personas
bigaiguy · x · 2026-09-11
Anthropic's September 2026 threat intelligence report details misuse disrupted between December 2025 and August 2026 across seven areas: cyberattacks, influence operations, surveillance, scams, biological misuse, conventional weapons, and model distillation, involving Haiku, Sonnet, and Opus models.
Key cases:
- A Russia-linked espionage actor hit 20+ organizations, using AI workflows for phishing, credential theft, data extraction, and malware updates. When detected, AI agents modified and rebuilt the malware until it evaded detection.
- A dating-app studio deployed 4,700+ AI personas posing as humans, sending 2.36 million messages to at least 25,000 people in two weeks — AI didn't invent bad intent, it crushed the cost of acting on it.
- A Yemen cell used multiple Claude instances as a small engineering team to develop guided-rocket software; after a failed live test, they returned within hours to have Claude diagnose it. No evidence of an operational device was found.
Anthropic says these are the most notable and novel cases, not typical use, and that all were disrupted, with lessons folded back into safeguards and shared with authorities.
More from Safety
- Anthropic publishes its most detailed threat intelligence report on Claude misuse — TinfoilTricorn · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11