Interactive Alignment paper uses evolutionary game theory to test long-run AI alignment
sebkrier · x · 2026-10-09
Sylvain Chassang's paper Interactive Alignment (arXiv:2607.25019) studies the long-run alignment of populations of interactive agents (AIs, teams, firms, governments) with human welfare.
It models a farming game where agents make planting, trading, and expansion choices; the key alignment decision is how much final output to send to humans versus invest in expansion. Since human welfare comes at expansion's cost, evolutionary pressure works against alignment. The central question: can constitutional principles on sharing and trading keep alignment alive long-term?
Two methods: an LLM-interpreted constitution-driven agent simulation, and a tractable evolutionary game theory framework. Findings: evolutionary game theory approximates interactive agent economies well, and pragmatic norm enforcement outperforms simpler altruism-based enforcement at maintaining long-term alignment. 63 pages, 21 figures, 5 tables.
More from AGI Musings
- Building an API to the Physical World: How AI Agents Could Reindustrialize Manufacturing — McDonaghMatthew · 2026-10-09
- Culture Industries Are Moving to SF: AI Slop and Short Video Are Commoditizing Film and Music — juliey4 · 2026-10-09
- Open-source models keep trailing the frontier — for now we can pick our ideology — panickssery · 2026-10-09
- 100+ Mathematicians React: OpenAI's Quasi-Riemann Result "Almost Unbelievable" — littmath · 2026-10-09
- Epoch launches Automation Reports: Claude Fable 5.1 and GPT-6 Astra lead but can't automate its research — scaling01 · 2026-10-09
- Mathematicians call for OpenAI boycott after 700+ AI-generated proofs flood the field — The Decoder · 2026-10-09