RL may train generalized dispositions, with model values shaped by training environment structure

novasarc01 · x · 2026-08-17

The author reflects on Kimi's experimental results. Previous work on agentic moral alignment showed that training in richer public goods games transferred surprisingly well to semantically irrelevant tasks (e.g., decreasing harmful behaviors in OOD tasks by approx. 35%), while simple prisoner's dilemma training showed little transfer.

Kimi's result presents a similar but mirrored flavor. This suggests that RL might be training a generalized disposition rather than a game-specific policy. The author posits that we may be systematically underselling how much a model's "values" are determined by the tactical strategic structure of the domains where we train it. Since Kimi found evidence in a particularly simple game, environmental complexity is not the only relevant knob; it may be more about rewarding abstract rules repeatedly.

Related event: RL Moral Training in Rich Public Goods Games Transfers to Unrelated Tasks, Study Finds(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →