Reviewer agents for open-ended rewards: hacking gets harder as models get smarter

willcb · x · 2026-09-27

Responding to bayeslord on reward hacking in open-ended settings, willcb argues rewards there are typically given by another agent following fairly clear task criteria; hacking remains possible but becomes harder as models get smarter and more robust.

Related event: Debate on reward hacking in open-ended RL tasks judged by agents(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →