davidad conjectures robust RL basins and self-DPO share one principle: dual model calls

davidad · x · 2026-09-04

AI safety researcher davidad responds to new alignment work with a conjecture: the discussed mechanism and his previously discussed self-DPO may be instances of a common principle — to get a model into a robust basin, the RL scorer must use the model at least twice independently, since basin shape requires nonlinearity. He also points to hyperproperties and the pRHL→eRHL upgrade as formal-methods reference points.

Related event: davidad conjectures cross-AI reward and self-DPO share one mechanism(2 posts)→

Original post →

More from Research

Research channel →