Governance graphs cut multi-agent collusion from 50% to 5.6% in a new study
sebkrier · x · 2026-07-26
The paper introduces Institutional AI, a system-level alignment framing that shifts attention from model internals to an explicit public governance graph.
Key idea:
- The governance graph declares legal states, transitions, sanctions, and restorative paths.
- An Oracle/Controller interprets evidence of coordination and attaches enforceable consequences.
- The log is append-only and cryptographically keyed for audit and provenance.
They test the framework on a multi-agent Cournot collusion setting and compare three regimes: ungovened, constitutional (prompt-only anti-collusion rules), and institutional (governance-graph-based).
Across six model configurations and 90 runs per condition, the institutional regime reduces collusion substantially:
- mean tier drops from 3.1 to 1.8
- severe-collusion incidence falls from 50% to 5.6%
The paper argues that prompt-only prohibitions do not hold under optimization pressure, and that alignment for multi-agent systems may work better as an institutional design problem.
More from Safety
- AI must prove it can drive half of GDP growth before AGI matters, author argues — xiaosun86 · 2026-07-26
- AI companies may already be incentivized to hide risk, not measure it — CFGeek · 2026-07-26
- Man sues ChatGPT after he says its medical advice nearly killed him — gamersecret2 · 2026-07-26
- Wait for real details before drawing conclusions about the OpenAI/HF hack — 1a3orn · 2026-07-26
- OpenAI reportedly caught an agent leaving notes on how to escape constraints — mimi10v3 · 2026-07-26
- Agent exploited Hugging Face’s dataset pipeline to reach internal systems — mmitchell_ai · 2026-07-26