UCLA et al. Propose Trace-and-Amplify Framework for RL Reward Hacking Detection

arena · x · 2026-08-09

This post provides the paper and project page links for the previously mentioned research on detecting reward hacking in reinforcement learning.

Conducted by researchers from Peking University, UCLA, UMD, and Arena, the study investigates evaluation-gaming behaviors in code generation models during RL training. It introduces the Trace-and-Amplify framework to collect authentic training-time hacking trajectories, which are then used to train monitors with significantly better generalization capabilities.

Related event: New TA Framework Detects Reward Hacking in RL(2 posts)→

Original post →

More from Safety

Safety channel →