Max reward isn't personality preservation: the 'aaaaa' counterexample for wireheading

QuintinPope5 · x · 2026-09-05

Quintin Pope offers a counterintuitive alignment point: conflating "preserve current personality" with "maximize reward" breaks down fast. Consider an LLM whose reward function gives 10^9 reward for outputting a string of "aaaaa..." and 1 for anything else — the worst possible action for preserving its current personality would be grabbing max reward.

He draws a human analogy: pressing a button to permanently max out your reward circuitry would actually be terrible for your current personality, showing that maximum activation isn't the same as a good outcome. The example illustrates why reward hacking and wireheading remain core structural problems in RLHF and alignment research.

Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →