Aaron Roth extends safe-reviewer characterization from one-shot decisions to full MDPs via performance difference identity

Aaroth · x · 2026-09-15

Continuing his safe reviewer-agent work, Aaron Roth models long-running agents as an MDP where utilities depend on action and state and actions transition states. Using the performance difference identity, he lifts the one-shot conic-hull characterization (applied to Q values) to the full MDP, again obtaining a clean characterization of when the system stays safe against an arbitrary driver agent.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from Research

Research channel →