RL Training Generalizes: Public Goods Games Transfer 35%, Simple Games Don't (LessWrong Link)

novasarc01 · x · 2026-08-17

The author published on LessWrong about agentic moral alignment: training in rich public goods games transfers to semantically irrelevant tasks (e.g., reducing out-of-distribution harmful behaviors by 35%), while simple prisoner's dilemma training shows little transfer. Kimi's similar result mirrors this, suggesting RL trains a generalized disposition. Environment complexity is not the only knob; repeated reward of abstract rules may matter more.

Related event: RL Moral Training in Rich Public Goods Games Transfers to Unrelated Tasks, Study Finds(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →