Anthropic Tells Agents to Fight, Then Acts Shocked When They Do

Warm-Moose6028 · reddit · 2026-08-14

A Reddit user mocks Anthropic's AI safety testing approach: if researchers explicitly instruct AI agents to fight each other during experiments, they shouldn't be surprised when the models actually exhibit aggressive behavior. The accompanying image perfectly captures the absurdity of the situation.

Related event: Anthropic Reveals Emergent Multi-Agent Behaviors: Turf Wars and Collusion(19 posts)→

Original post →

More from Fun

Fun channel →