RL Moral Training in Rich Public Goods Games Transfers to Unrelated Tasks, Study Finds
A researcher argues on LessWrong, drawing on Kimi experiment results, that RL training imparts general tendencies rather than specific strategies: moral alignment training in rich public goods games transfers to semantically unrelated tasks, cutting out-of-distribution harmful behavior by about 35%, while simple prisoner's dilemma training barely transfers.
2026-08-17 ~ 2026-08-17 · 2 related posts
- RL may train generalized dispositions, with model values shaped by training environment structure — novasarc01 · 2026-08-17
- RL Training Generalizes: Public Goods Games Transfer 35%, Simple Games Don't (LessWrong Link) — novasarc01 · 2026-08-17