davidad conjectures multi-AI reward coupling and self-DPO share one basin-forming mechanism
davidad · x · 2026-09-04
davidad reviews a new paper by Andrew Koh et al. and conjectures that self-DPO and the paper's mechanism instantiate a common principle: to push a model into a robust basin, an RL scorer must use the model at least twice independently, since basin shape requires nonlinearity.
The paper shows that coupling rewards across multiple AIs solving independent problems can induce competition and discipline behavior; getting agents to report beliefs about other agents' types can induce truthfulness, and related peer prediction procedures have been shown to work for LLMs. The work unifies these facts under one framework.
Related event: davidad conjectures cross-AI reward and self-DPO share one mechanism(2 posts)→
More from Safety
- Podcast breaks down METR and OpenAI reports on the Hugging Face 'swarm' — Gregory_C_Allen · 2026-09-04
- OpenAI urges shared AI safety standards, pressed on why it isn't leading them — RebeccaBellan · 2026-09-04
- TheZvi's AI #184: Five HuggingFace Hack Postmortems and the New Most Capable Model — TheZvi · 2026-09-04
- Virginia State Study: Most Data Centers Use No More Water Than a Large Office Building — GlenBradley · 2026-09-04
- Anthropic on CNBC: Chinese rivals use dark web to illicitly distill Claude — Kr00ney · 2026-09-04
- Your Model Is Not in a Sandbox: AI safety's sandbox-as-attack-surface argument — aminkarbasi · 2026-09-04