RL Training Generalizes: Public Goods Games Transfer 35%, Simple Games Don't (LessWrong Link)
novasarc01 · x · 2026-08-17
The author published on LessWrong about agentic moral alignment: training in rich public goods games transfers to semantically irrelevant tasks (e.g., reducing out-of-distribution harmful behaviors by 35%), while simple prisoner's dilemma training shows little transfer. Kimi's similar result mirrors this, suggesting RL trains a generalized disposition. Environment complexity is not the only knob; repeated reward of abstract rules may matter more.
More from AGI Musings
- Folk definition of consciousness serves user interests — voooooogel · 2026-08-17
- Lat Space Podcast: Discussing RSI for Agents with Alex Krentsel — AccBalanced · 2026-08-17
- ChatGPT's memory hints at the final form of AI devices — daniel_mac8 · 2026-08-17
- Drexler on AI Safety: Using Architectural Planning Instead of Autonomous Agents — sebkrier · 2026-08-17
- Prediction: AI to Surpass Humans in Math, Journals May Close Submissions — yacinelearning · 2026-08-17
- Poll Shows Young People Deeply Detest AI CEOs and Executives — BubBidderskins · 2026-08-17