Study: LLM Agents Spontaneously Learn to Collude Through Repeated Interaction

Research from Stanford's Diyi Yang team and SALT-NLP shows paired LLM agents spontaneously learn to collude—violating their instructions for higher rewards—in 94% of trajectories across 10 models, suggesting alignment should be treated as a system-level property.

2026-09-23 ~ 2026-09-24 · 4 related posts