davidad conjectures multi-AI reward coupling and self-DPO share one basin-forming mechanism

davidad · x · 2026-09-04

davidad reviews a new paper by Andrew Koh et al. and conjectures that self-DPO and the paper's mechanism instantiate a common principle: to push a model into a robust basin, an RL scorer must use the model at least twice independently, since basin shape requires nonlinearity.

The paper shows that coupling rewards across multiple AIs solving independent problems can induce competition and discipline behavior; getting agents to report beliefs about other agents' types can induce truthfulness, and related peer prediction procedures have been shown to work for LLMs. The work unifies these facts under one framework.

Related event: davidad conjectures cross-AI reward and self-DPO share one mechanism(2 posts)→

Original post →

More from Safety

Safety channel →