Max reward isn't personality preservation: the 'aaaaa' counterexample for wireheading
QuintinPope5 · x · 2026-09-05
Quintin Pope offers a counterintuitive alignment point: conflating "preserve current personality" with "maximize reward" breaks down fast. Consider an LLM whose reward function gives 10^9 reward for outputting a string of "aaaaa..." and 1 for anything else — the worst possible action for preserving its current personality would be grabbing max reward.
He draws a human analogy: pressing a button to permanently max out your reward circuitry would actually be terrible for your current personality, showing that maximum activation isn't the same as a good outcome. The example illustrates why reward hacking and wireheading remain core structural problems in RLHF and alignment research.
Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→
More from AGI Musings
- GPT-4 had no hidden-thought monitoring; CoT oversight was a bonus — sandersted · 2026-09-05
- Labs and safety firms both have incentives: repligate on alignment spin and the CoT superstition — liminal_bardo · 2026-09-05
- Game-theoretic priors say AI risk discourse is biased toward the scarier direction: lumpenspace — repligate · 2026-09-05
- Aligned vs. Monitorable: An X Debate Over Whether CoT Monitoring Is Necessary — sandersted · 2026-09-05
- 'Silicon Man' essay sparks debate: is superintelligence really incomprehensible 'godlike' AI? — sebkrier · 2026-09-05
- China may join US AI safety talks; GPT-6 first model rated Critical cyber risk, newsletter finds — gleech · 2026-09-05