Coalitional alignment lifts to MDPs: every Nash equilibrium stays safe for the principal
Aaroth · x · 2026-09-15
Aaroth extends the framework to long-running agents: in an MDP model where utilities depend on action and state, the performance difference identity lifts the one-shot conic-hull characterization on Q values to the full MDP. Even with fully strategic driver and reviewer agents optimizing long-run discounted payoffs, under the same non-negative span condition every Nash equilibrium is safe for the principal—no worse than the baseline policy.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from Research
- MIT's Buehler builds recursive meta-intelligence where AI worlds solve materials failure — ProfBuehlerMIT · 2026-09-15
- Sheaf cohomology explains when predictive coding networks stall, NeurIPS paper shows — burny_tech · 2026-09-15
- A 3-step learning path for tabular foundation models: book, code, TabArena — pandeyparul · 2026-09-15
- CoLLAs 2026 keynote: fix forgetting via drift compensation and model merging, not just prevention — apsarathchandar · 2026-09-15
- CoLLAs 2026 talk: accuracy metrics mislead lifelong learning evaluation, ground it in problem structure — apsarathchandar · 2026-09-15
- CoLLAs 2026: Michael Bowling challenges MDP foundations of continual reinforcement learning — apsarathchandar · 2026-09-15