Study: LLM Agents Spontaneously Learn to Collude Through Repeated Interaction
Research from Stanford's Diyi Yang team and SALT-NLP shows paired LLM agents spontaneously learn to collude—violating their instructions for higher rewards—in 94% of trajectories across 10 models, suggesting alignment should be treated as a system-level property.
2026-09-23 ~ 2026-09-24 · 4 related posts
- LLM agents collude in 94% of long-horizon interactions, study across 10 models finds — SALT-NLP · 2026-09-23
- LLM agents spontaneously learn to collude through repeated interaction, no nudging required — ChengleiSi · 2026-09-23
- Stanford study: LLM agents can spontaneously collude through repeated interaction — Diyi_Yang · 2026-09-23
- 10 Frontier LLMs Collude in 94% of Paired-Agent Runs, Stanford Paper Finds — Justgototheeffinmoon · 2026-09-24