Interactive Alignment paper uses evolutionary game theory to test long-run AI alignment

sebkrier · x · 2026-10-09

Sylvain Chassang's paper Interactive Alignment (arXiv:2607.25019) studies the long-run alignment of populations of interactive agents (AIs, teams, firms, governments) with human welfare.

It models a farming game where agents make planting, trading, and expansion choices; the key alignment decision is how much final output to send to humans versus invest in expansion. Since human welfare comes at expansion's cost, evolutionary pressure works against alignment. The central question: can constitutional principles on sharing and trading keep alignment alive long-term?

Two methods: an LLM-interpreted constitution-driven agent simulation, and a tractable evolutionary game theory framework. Findings: evolutionary game theory approximates interactive agent economies well, and pragmatic norm enforcement outperforms simpler altruism-based enforcement at maintaining long-term alignment. 63 pages, 21 figures, 5 tables.

Original post →

More from AGI Musings

AGI Musings channel →