Real reward models show safety-completeness tradeoffs from coalitional alignment

Aaroth · x · 2026-09-15

Aaroth notes the coalitional alignment condition is not guaranteed to hold, but preliminary experiments on real reward models reveal non-trivial safety/completeness tradeoffs, driven by coalitional rather than individual alignment.

Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→

Original post →

More from Research

Research channel →