Extreme reward functions push LLMs away from their current behavior

jessi_cata · x · 2026-09-05

A technical exchange on RL reward design: the author assumes variance control as the RL baseline but notes a run can still "luck into" a behavior that then gets reinforced. Counterexample: if a reward function gives 10^9 reward for outputting "aaaaa..." and 1 otherwise, the worst thing an LLM can do to preserve its current personality is maximize reward — extreme reward settings systematically push models away from their existing behavior.

Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→

Original post →

More from Research

Research channel →