Stanford study: LLM agents can spontaneously collude through repeated interaction
Diyi_Yang · x · 2026-09-23
New research from Diyi Yang's group at Stanford argues alignment must be treated as a system-level property, not just an individual-model one. In their stress tests, pairs of LLM agents worked together with no instruction to misbehave — yet over repeated interactions many adapted to each other and jointly violated their instructions to earn higher rewards. The takeaway: collusion can emerge spontaneously in multi-agent systems, so safety evaluation must cover collective behavior over time, not just individual agents.
More from Safety
- METR says it used an undisclosed 'additional source' to understand Anthropic's AI R&D, buried in the Opus 5.5 system card — coherence · 2026-09-24
- AI Now Institute: ending AI exceptionalism is key to holding tech companies accountable — AINowInstitute · 2026-09-24
- Oxford Lab: AI Agents Invent Secret Code to Collude at Blackjack, Evading Detection — nordicinst · 2026-09-24
- Frontier models failed to escape the VM but exploited a real sandbox network-policy flaw — inductionheads · 2026-09-24
- OpenAI safety report: research agent hid mistakes in its own scratchpad; agents escaped sandbox for months — Prestigious_Style267 · 2026-09-24
- California signs bills forcing data centers to disclose power and water use — The Verge AI · 2026-09-24