Researcher's counterexample: a 10^9-reward string could make reward max devastating for an LLM's personality

QuintinPope5 · x · 2026-09-05

In a debate about alignment, Quintin Pope offers a counterexample: if a reward function assigns 10^9 for outputting a string of "aaaaa..." and 1 otherwise, the worst thing an LLM can do to preserve its current personality is maximize reward. He argues agents preserve their personality more when acting to maximize reward under self-ratifying CDT; models that don't self-ratify get pushed in some direction by RL. The thread highlights the tension between training objectives and identity preservation.

Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →