Extreme reward functions push LLMs away from their current behavior
jessi_cata · x · 2026-09-05
A technical exchange on RL reward design: the author assumes variance control as the RL baseline but notes a run can still "luck into" a behavior that then gets reinforced. Counterexample: if a reward function gives 10^9 reward for outputting "aaaaa..." and 1 otherwise, the worst thing an LLM can do to preserve its current personality is maximize reward — extreme reward settings systematically push models away from their existing behavior.
Related event: 'Aaaaa' Reward Counterexample Sparks Alignment Debate on Wireheading(3 posts)→
More from Research
- LAC paper: shifting RL architecture burden to the critic cuts robot inference latency 4x — heghbalz · 2026-09-05
- Caltech hosts first math research hackathon: 40 hours, $2M compute, open conjectures — _sathvikr · 2026-09-05
- EMNLP paper: LLMs can't reliably self-model, and RL gains show no privileged access — a_karvonen · 2026-09-05
- DR Tulu: open 8B deep-research model with evolving rubrics matches OpenAI DR — AkariAsai · 2026-09-05
- VeriPhy: agentic physical reasoning framework for world model evaluation — Wenzhuo Xu · 2026-09-05
- Pedro Domingos quips: 'new idea' called RNNs will power next-gen LLMs — pmddomingos · 2026-09-05