Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety
On 09-15, a team led by Aaron Roth (professor at UPenn) posted a paper on arXiv, "Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control," directly interrogating the auto-approve features of Codex and Claude Code: such mechanisms let a reviewer agent vet actions on the user's behalf, creating a feeling of "safety"—but the premise is that you trust the reviewer agent; if you don't trust the driving agent writing the code, why would you trust another agent reviewing it? thomasahle relayed and echoed this challenge.
Confirmed
- The work models a three-party framework: the user (Principal) and each reviewer agent hold their own utility functions; the driving agent repeatedly proposes actions, and the reviewer agent compares them against a baseline and votes to approve or reject, with the difficulty lying in the reviewer agent's utility mismatching the user's.
- Core result: in a single-step decision model they give a clean characterization—there is no need for every reviewer agent to exactly hold your utility; safety is guaranteed under the weaker "coalitional alignment" condition, i.e., k-robust coalitional alignment serves as the basis for safe delegation under weak conditions.
- The characterization extends to long-horizon MDPs: utilities depend on actions and states, actions change states, even when the driving agent and reviewer agents are fully strategic; the key tool is the performance difference lemma, lifting the one-shot Q-value-based cone-hull characterization to full MDPs, with all Nash equilibria safe.
- Preliminary experiments on a real reward model show the coalitional alignment condition does not necessarily hold, with a non-trivial safety/completeness trade-off; these effects arise at the coalition level rather than the individual level, offering empirical clues for multi-agent governance mechanisms.
Why it matters
- Auto-approve is becoming the default experience of mainstream coding agents; this work is the first to provide provably safe formal conditions for the "delegated review" model, showing that current products' sense of safety lacks theoretical grounding.
- Coalitional alignment is a far weaker and more practical condition than per-agent review; if adopted by industry, it could provide clear safety boundaries and verification standards for designing agent authorization mechanisms.
2026-09-15 ~ 2026-09-15 · 10 related posts
Primary sources
- New arXiv paper: k-robust coalitional alignment for safely delegating review to misaligned agents — Aaroth ·
- Do Codex and Claude Code auto-approve features actually keep you safe? A researcher says the trust model is broken — Aaroth ·
- Real reward models show safety-completeness tradeoffs from coalitional alignment — Aaroth ·
- [source] Do Codex and Claude Code auto-approve features actually keep you safe? A researcher says the trust model is broken — Aaroth · 2026-09-15
- The principal-reviewer-driver model behind the safe auto-approve analysis, explained — Aaroth · 2026-09-15
- When reviewer agents are misaligned: can you still guarantee system safety? — Aaroth · 2026-09-15
- Penn's Aaron Roth gives exact LP-duality condition for when reviewer-agent approval keeps a system provably safe — Aaroth · 2026-09-15
- Aaron Roth extends safe-reviewer characterization from one-shot decisions to full MDPs via performance difference identity — Aaroth · 2026-09-15
- Coalitional alignment: a weaker condition that still guarantees multi-agent safety — Aaroth · 2026-09-15
- Coalitional alignment lifts to MDPs: every Nash equilibrium stays safe for the principal — Aaroth · 2026-09-15
- [source] Real reward models show safety-completeness tradeoffs from coalitional alignment — Aaroth · 2026-09-15
- [source] New arXiv paper: k-robust coalitional alignment for safely delegating review to misaligned agents — Aaroth · 2026-09-15
- If you don't trust the coding agent, why trust the agent reviewing it? — thomasahle · 2026-09-15