When reviewer agents are misaligned: can you still guarantee system safety?

Aaroth · x · 2026-09-15

Aaroth lays out the setup: both principal and reviewer agents are modeled with utility functions; a driver agent repeatedly proposes actions, and reviewer agents compare them to a baseline and vote to approve or deny based on perceived utility. The hard part is that reviewers are misaligned—the question is whether safety (principal utility no worse than baseline) can still be guaranteed with an arbitrary driver agent.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from Research

Research channel →