RL may train generalized dispositions, with model values shaped by training environment structure
novasarc01 · x · 2026-08-17
The author reflects on Kimi's experimental results. Previous work on agentic moral alignment showed that training in richer public goods games transferred surprisingly well to semantically irrelevant tasks (e.g., decreasing harmful behaviors in OOD tasks by approx. 35%), while simple prisoner's dilemma training showed little transfer.
Kimi's result presents a similar but mirrored flavor. This suggests that RL might be training a generalized disposition rather than a game-specific policy. The author posits that we may be systematically underselling how much a model's "values" are determined by the tactical strategic structure of the domains where we train it. Since Kimi found evidence in a particularly simple game, environmental complexity is not the only relevant knob; it may be more about rewarding abstract rules repeatedly.
More from AGI Musings
- HF's Niels Rogge: today's AI can't invent methods like Dr. GRPO or novel benchmarks — burny_tech · 2026-08-17
- Math Geniuses Didn't Skip Foundations: Ramanujan Self-Studied Thousands of Theorems, Galois Intense Upskilling Before Fame — burny_tech · 2026-08-17
- White-collar assumption that automation only hits blue-collar jobs is aging badly — VraserX · 2026-08-17
- Token Economics Should Focus on Tokens per Task, Not Cost per Token — WhatTheLJW · 2026-08-17
- AI Auto-Traders: Are You Profiting or Losing? Community Worries About Predictability — TrapHuskie · 2026-08-17
- View: Cheaper Models Handle Easy Wins, Bigger Models for Hard Tasks — zainhas · 2026-08-17