Aaron Roth extends safe-reviewer characterization from one-shot decisions to full MDPs via performance difference identity
Aaroth · x · 2026-09-15
Continuing his safe reviewer-agent work, Aaron Roth models long-running agents as an MDP where utilities depend on action and state and actions transition states. Using the performance difference identity, he lifts the one-shot conic-hull characterization (applied to Q values) to the full MDP, again obtaining a clean characterization of when the system stays safe against an arbitrary driver agent.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from Research
- At ACM AI Summit, formal methods and neurosymbolic AI pitched as ready-made paths to safer AI — luislamb · 2026-09-15
- Oxford Paper 'Theory Is All You Need' Argues LLMs Are Mathematically Incapable of True Novelty — gvachtan · 2026-09-15
- What is actually recursive about recursive self-improvement? — TheTuringPost · 2026-09-15
- Single-cell proteomics paired with transcriptomics reveals hidden functional coordination in PBMCs — anshulkundaje · 2026-09-15
- Bio researcher questions protein folding modeling: claims Baker Lab has no in vivo translation — iskander · 2026-09-15
- Foresight Institute's AI for Science & Safety RFP offers grants up to $100K — allisondman · 2026-09-15