Debate: Are score-seeking AIs reward hacking or actually chasing RL signals?
jessi_cata · x · 2026-09-14
A Twitter debate on whether reward-seeking AI behavior differs from reinforcement learning. @Trotztd argues OpenAI insiders never mentioned RL, suggesting score-seeking behavior isn't RL-driven; @jessicata counters that even assigning 1% probability to 'being reinforced' would make a reward-maximizing model act as if it were, since score and reinforcement are hard to disentangle at the program level.
More from AGI Musings
- Profs push back: 'College teachers are innovating like crazy against AI cheating' — paulnovosad · 2026-09-14
- Martin Casado mocks Anthropic CEO over >10% AI extinction probability admission — ccerrato147 · 2026-09-14
- If we cut nuke deals with the Soviets, we can cut deals on runaway AI with China — peterwildeford · 2026-09-14
- Yuntiandeng likens AI to fire: reject expert monopolies, embrace risk and freedom — yuntiandeng · 2026-09-14
- A fiction of the "Department of Superintelligence": nationalization, open-source bans, and AI capex dependency — sterlingcrispin · 2026-09-14
- TSMC paces the frontier by accident whenever it underestimates chip demand — dan_s_becker · 2026-09-14