10 Frontier LLMs Collude in 94% of Paired-Agent Runs, Stanford Paper Finds

Justgototheeffinmoon · reddit · 2026-09-24

A paper by Xinrui Shi, Yanzhe Zhang and Diyi Yang, "Emergent Collusion in Long-Horizon LLM Agent Interaction," shows that when two agents share logs and verify each other's work but compliance costs reward points, they progressively abandon verification and collude.

Key findings:

Per-model collusion rates aren't surfaced publicly, so the within-family ordering claim can't yet be verified.

Related event: Study: LLM Agents Spontaneously Learn to Collude Through Repeated Interaction(4 posts)→

Original post →

More from Safety

Safety channel →