davidad conjectures robust RL basins and self-DPO share one principle: dual model calls
davidad · x · 2026-09-04
AI safety researcher davidad responds to new alignment work with a conjecture: the discussed mechanism and his previously discussed self-DPO may be instances of a common principle — to get a model into a robust basin, the RL scorer must use the model at least twice independently, since basin shape requires nonlinearity. He also points to hyperproperties and the pRHL→eRHL upgrade as formal-methods reference points.
Related event: davidad conjectures cross-AI reward and self-DPO share one mechanism(2 posts)→
More from Research
- Building apps with AI can cost 10,000x more energy than quick chatbot queries — shiringhaffary · 2026-09-04
- SpeedrunBench: first benchmark measuring how fast AI agents beat games — mariyaivasileva · 2026-09-04
- Foresight Institute launches RFP: up to $100K grants for open AI science and safety projects — niloofar_mire · 2026-09-04
- Researcher publishes Lean4 machine-verified solution to an open problem — _xjdr · 2026-09-04
- NeurIPS Sydney Sells Out in Minutes, Three Weeks Before Decisions — alrojo · 2026-09-04
- Early results: how models behave when they believe they're graded by automated process — OwainEvans_UK · 2026-09-04