Monitor RL rollout actions to catch reward hacking, not persona drift, argues voooooogel
voooooogel · x · 2026-09-06
@voooooogel offers a methodological takeaway in the same alignment thread: to robustly detect reward hacking, monitor the RL rollout actions directly. Detecting changes via weird persona generalization was even less robust to deception than CoT monitoring, he argues, keeping with his skepticism of persona-based detection approaches.
More from Safety
- Creator discloses sponsor paid for AI-dangers videos, sparking disclosure debate — jessi_cata · 2026-09-06
- Creators Reveal Paid AI Doomer Outreach; Guardian Article's 6 Participants All Donor-Funded — beffjezos · 2026-09-06
- OpenAI researcher rebuts ex-NCSC chief: AI agents in incidents weren't following orders — Miles_Brundage · 2026-09-06
- Sanders' superintelligence ban bill slammed: 20-year prison terms borrowed from nuclear statute — r0ck3t23 · 2026-09-06
- Opinion: calling AI a 'rogue agent' is corporate liability-dodging — ai · 2026-09-06
- An algorithm off switch isn't enough: Zoe Daniel calls for a digital duty of care on tech giants — nordicinst · 2026-09-06