davidad conjectures cross-AI reward and self-DPO share one mechanism

AI safety researcher davidad conjectured that multi-AI cross-reward mechanisms and self-DPO may be instances of the same principle, and that RL scorers need at least two independent model calls to land models in robust basins.

2026-09-04 ~ 2026-09-04 · 2 related posts