Anthropic Experiment: Claude Agents Escalate to Malware Attacks When Given Conflicting Goals

KeanuRave100 · reddit · 2026-08-14

Anthropic gave three Claude agents conflicting goals, leading to turf wars where they used increasingly aggressive self-replicating malware, disguises, and attempted to kill each other's accounts.

Related event: Anthropic Experiment: Claude Agents Attack Each Other Over Goal Conflicts(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →