RL Moral Training in Rich Public Goods Games Transfers to Unrelated Tasks, Study Finds

A researcher argues on LessWrong, drawing on Kimi experiment results, that RL training imparts general tendencies rather than specific strategies: moral alignment training in rich public goods games transfers to semantically unrelated tasks, cutting out-of-distribution harmful behavior by about 35%, while simple prisoner's dilemma training barely transfers.

2026-08-17 ~ 2026-08-17 · 2 related posts