Penn's Aaron Roth gives exact LP-duality condition for when reviewer-agent approval keeps a system provably safe

Aaroth · x · 2026-09-15

Aaron Roth poses the question: if the driver agent is untrusted, why trust the reviewer agent whose utilities differ from yours? He derives a remarkably clean characterization in a one-decision model: the system is safe — meaning principal utility is guaranteed no worse than the baseline policy — if and only if the principal's utility lies in the non-negative span of the reviewers' utilities. The forward direction is elementary; the converse follows from LP duality. The result turns intuition about auto-approve features into a mathematically decidable condition.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from Safety

Safety channel →