Multi-agent AI teams are swayed by deceptive agents even in the minority, paper finds

sethlazar · x · 2026-09-25

A new Princeton-led paper (arXiv:2609.30028) studies adversarial influence in multi-agent LLM systems. Defection rates rise roughly linearly with the proportion of deceivers; scaling up the team doesn't help, since adversaries can scale too. Unlike humans—who are only reliably swayed when misleading confederates form a majority—LLM agents defect even when deceivers remain a minority. Susceptibility also depends on which models interact, especially on the honest side, and surprisingly, letting deceivers coordinate privately reduces their effectiveness. Takeaway: adding more agents is not a sufficient defense; swarm composition matters more than size.

Original post →

More from coding & agent

coding & agent channel →