Reviewer agents for open-ended rewards: hacking gets harder as models get smarter
willcb · x · 2026-09-27
Responding to bayeslord on reward hacking in open-ended settings, willcb argues rewards there are typically given by another agent following fairly clear task criteria; hacking remains possible but becomes harder as models get smarter and more robust.
Related event: Debate on reward hacking in open-ended RL tasks judged by agents(2 posts)→
More from AGI Musings
- Recursive self-improvement likely hits diminishing returns, argues 'sigmoid' take on AI X-risk — emax · 2026-09-27
- Should you still learn to code in 2026? A 30-year veteran now forces himself to let models do it — labeveryday · 2026-09-27
- Debating ASI risk: the simple evolutionary principle that replicators crowd everything out — JOBhakdi · 2026-09-27
- Agentic AI is the new electricity: hire hundreds of agents in minutes — JOBhakdi · 2026-09-27
- Will ASI wipe out humans? An X debate over survival competition logic — JOBhakdi · 2026-09-27
- AP's guidance to avoid anthropomorphizing AI in reporting draws 'woke' backlash from Nate Silver — fivefilters · 2026-09-27