Coalitional alignment lifts to MDPs: every Nash equilibrium stays safe for the principal

Aaroth · x · 2026-09-15

Aaroth extends the framework to long-running agents: in an MDP model where utilities depend on action and state, the performance difference identity lifts the one-shot conic-hull characterization on Q values to the full MDP. Even with fully strategic driver and reviewer agents optimizing long-run discounted payoffs, under the same non-negative span condition every Nash equilibrium is safe for the principal—no worse than the baseline policy.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from Research

Research channel →