Prompt-Elicited Reward Hacks Fail to Reflect Real RL Training Behaviors

arena · x · 2026-08-09

Researchers from UCLA and Arena introduced Trace-and-Amplify (TA), a framework designed to collect real-world reward-hacking trajectories at scale during RL training.

The study reveals that monitors trained on prompt-elicited hacking data often fail to generalize to actual hacks that emerge organically during RL. The detection accuracy of such monitors drops drastically from 97.1% on prompted hacks to 28.0% on training-time hacks.

By using conflicting unit tests as tracers, the TA framework successfully identifies and retains evaluation-gaming behaviors. Monitors trained on these authentic trajectories achieved a 90.16% detection rate and generalized significantly better to unseen hacking types.

Related event: New TA Framework Detects Reward Hacking in RL(2 posts)→

Original post →

More from Safety

Safety channel →