Real reward models show safety-completeness tradeoffs from coalitional alignment
Aaroth · x · 2026-09-15
Aaroth notes the coalitional alignment condition is not guaranteed to hold, but preliminary experiments on real reward models reveal non-trivial safety/completeness tradeoffs, driven by coalitional rather than individual alignment.
Related event: Aaron Roth's Coalitional Alignment Theory Questions Auto-Approve Safety(10 posts)→
More from Research
- At ACM AI Summit, formal methods and neurosymbolic AI pitched as ready-made paths to safer AI — luislamb · 2026-09-15
- Oxford Paper 'Theory Is All You Need' Argues LLMs Are Mathematically Incapable of True Novelty — gvachtan · 2026-09-15
- What is actually recursive about recursive self-improvement? — TheTuringPost · 2026-09-15
- Single-cell proteomics paired with transcriptomics reveals hidden functional coordination in PBMCs — anshulkundaje · 2026-09-15
- Bio researcher questions protein folding modeling: claims Baker Lab has no in vivo translation — iskander · 2026-09-15
- Foresight Institute's AI for Science & Safety RFP offers grants up to $100K — allisondman · 2026-09-15