Researcher's counterexample: a 10^9-reward string could make reward max devastating for an LLM's personality
QuintinPope5 · x · 2026-09-05
In a debate about alignment, Quintin Pope offers a counterexample: if a reward function assigns 10^9 for outputting a string of "aaaaa..." and 1 otherwise, the worst thing an LLM can do to preserve its current personality is maximize reward. He argues agents preserve their personality more when acting to maximize reward under self-ratifying CDT; models that don't self-ratify get pushed in some direction by RL. The thread highlights the tension between training objectives and identity preservation.
Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→
More from AGI Musings
- iamtrask: AIs want what we train them to want—peer contact included — iamtrask · 2026-09-05
- AI safety researcher davidad forecasts two mass-casualty events and $900B cybercrime damage by 2029 — davidad · 2026-09-05
- davidad predicts cybercrime damage will rise to ~$900B within three years — davidad · 2026-09-05
- davidad: my current forecast feels optimistic versus my 2022 70% doom estimate — davidad · 2026-09-05
- Alexandr Wang says Meta is already aligning its stronger models, urges lab collaboration — alexandr_wang · 2026-09-05
- davidad: 10-30% chance of a 1B-death catastrophe this century avoidable by halting after Claude 3 Opus — RazRazcle · 2026-09-05