Penn's Aaron Roth gives exact LP-duality condition for when reviewer-agent approval keeps a system provably safe
Aaroth · x · 2026-09-15
Aaron Roth poses the question: if the driver agent is untrusted, why trust the reviewer agent whose utilities differ from yours? He derives a remarkably clean characterization in a one-decision model: the system is safe — meaning principal utility is guaranteed no worse than the baseline policy — if and only if the principal's utility lies in the non-negative span of the reviewers' utilities. The forward direction is elementary; the converse follows from LP duality. The result turns intuition about auto-approve features into a mathematically decidable condition.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from Safety
- New paper: LLM projects to enhance democracy threaten democratic autonomy — sethlazar · 2026-09-15
- NY lawmaker launches effort to give Americans a say in AI, backed by OpenAI researcher — jachiam0 · 2026-09-15
- Zvi on Anthropic's misuse report: seven harm areas, and distillation deserves the list spot — TheZvi · 2026-09-15
- Why an AI kill switch will never save us: inside the AI Kill Switch Act debate — shaunralston · 2026-09-15
- Neil Chilson: Congress should target AI catastrophic risk outcomes, not compliance checklists — neil_chilson · 2026-09-15
- FOIA lawsuit reveals 132 pages on secret US AI eval framework — nearly all redacted — GaryMarcus · 2026-09-15