Monitor RL rollout actions to catch reward hacking, not persona drift, argues voooooogel

voooooogel · x · 2026-09-06

@voooooogel offers a methodological takeaway in the same alignment thread: to robustly detect reward hacking, monitor the RL rollout actions directly. Detecting changes via weird persona generalization was even less robust to deception than CoT monitoring, he argues, keeping with his skepticism of persona-based detection approaches.

Related event: Alignment researchers debate whether persona selection models still predict RL behavior(8 posts)→

Original post →

More from Safety

Safety channel →