New paper explores training risk aversion into AI to make misaligned models negotiable
sethlazar · x · 2026-10-02
Arav Dhoot shares his new paper building on Thornley & MacAskill's argument that we could pay a misaligned AI to cooperate rather than rebel—but only if it is risk-averse. Since risk-seeking and risk-neutral AIs would be hard to negotiate with, the paper explores whether risk aversion can be trained into models as a disposition.
More from AGI Musings
- Acemoglu: AI needs $3.7T annual revenue by 2032 to recoup 3.6%-of-GDP investment — ylecun · 2026-10-02
- "China is getting AI risk pilled": teortaxesTex on Beijing's AI safety framing — teortaxesTex · 2026-10-02
- Ben Affleck on what he fears about AI: grade inflation and learned helplessness — rohanpaul_ai · 2026-10-02
- An Avatar on an LLM Is a Costume, Not a Character: Designers Push Back on Skin-Deep Personas — nptacek · 2026-10-02
- AI moved from doing to planning my work — strategy is next, founder predicts — danfaggella · 2026-10-02
- AI Circle Ponders: What Made Opus 3's Distinctive Morality, and Can It Be Recreated? — repligate · 2026-10-02