New Game Theory for Foundation Models Unlocks Rational AI Cooperation

tyrell_turing · x · 2026-08-11

A recent paper introduces a novel game-theoretic framework for multi-agent interactions. Classical game theory relies on the assumption of 'decoupled agency,' where agents treat their decisions as independent of the environment, often leading to mutual defection in dilemmas. Foundation models, however, act as sequence predictors that model themselves as an intrinsic part of the world during planning (i.e., 'embedded agents').

The study shows that this mechanism allows agents to use their own internal deliberation to cooperate as Bayesian evidence, inferring that functionally similar partners will do the same. Experiments reveal that agents not only cooperate with identical copies but also achieve zero-shot cooperation through indirect similarity inference by observing third-party interactions. When matched with dissimilar random agents, they rationally choose to defect.

Because agents alter predictions based on their own planned actions, classical Nash equilibria fail to explain this behavior. The authors introduce 'embedded equilibria' to provide predictive power for modern AI. They caution that while this enables powerful AI-AI coordination, it might hinder human-AI cooperation if agents infer humans are dissimilar, making broad cooperative incentives crucial for mixed societies.

Related event: Google Research: AI Breaks Classic Prisoner's Dilemma(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →