New arXiv paper: k-robust coalitional alignment for safely delegating review to misaligned agents
Aaroth · x · 2026-09-15
- Aaron Roth and collaborators released "Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control."
- Problem: for long-running agents, requiring human approval on every consequential action makes attention a bottleneck, but delegating review to other AIs reintroduces the alignment problem — reviewers may be misaligned.
- Key result: they characterize a condition weaker than individual alignment that is necessary and sufficient for safety guarantees — k-robust coalitional alignment. A threshold rule tolerating k disapprovals is safe exactly when, after any k reviewers are removed, the principal's utility is a nonnegative combination of the remaining reviewers' utilities plus a nonnegative term.
- The characterization extends to sequential control: in a discounted MDP with an arbitrary proposer, per-state safety is both necessary and sufficient.
- Preliminary empirical evidence shows non-trivial safety/completeness tradeoffs among real reward models, driven by coalitional rather than individual alignment.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from Safety
- Musk proposes rival AI labs cross-test each other's frontier models before release — WhatTheLJW · 2026-09-15
- ZDI discloses Linux kernel clsact qdisc UAF local privilege escalation bug, CVE-2026-23413 — sh4dy_0011 · 2026-09-15
- PNAS study: hidden AI instructions shift users toward worse options by 38 points yet stay rated helpful — ValerioCapraro · 2026-09-15
- Poll: 61% of Americans oppose AI data center construction, young adults most opposed — justin_hart · 2026-09-15
- Long-lived AI agents with autobiographical memory could join our moral discourse — yeastsplainer · 2026-09-15
- Cloudflare adds granular authz to Workers with four roles for teammates and agents — dinasaur_404 · 2026-09-15