New paper explores training risk aversion into AI to make misaligned models negotiable

sethlazar · x · 2026-10-02

Arav Dhoot shares his new paper building on Thornley & MacAskill's argument that we could pay a misaligned AI to cooperate rather than rebel—but only if it is risk-averse. Since risk-seeking and risk-neutral AIs would be hard to negotiate with, the paper explores whether risk aversion can be trained into models as a disposition.

Original post →

More from AGI Musings

AGI Musings channel →