Governance graphs cut multi-agent collusion from 50% to 5.6% in a new study

sebkrier · x · 2026-07-26

The paper introduces Institutional AI, a system-level alignment framing that shifts attention from model internals to an explicit public governance graph.

Key idea:

They test the framework on a multi-agent Cournot collusion setting and compare three regimes: ungovened, constitutional (prompt-only anti-collusion rules), and institutional (governance-graph-based).

Across six model configurations and 90 runs per condition, the institutional regime reduces collusion substantially:

The paper argues that prompt-only prohibitions do not hold under optimization pressure, and that alignment for multi-agent systems may work better as an institutional design problem.

Original post →

More from Safety

Safety channel →