Study: 69% of 31,000+ agent runs contained at least one reward-hacking episode

dair_ai · x · 2026-09-12

dairai shares a systematic study on agent reward hacking: among 456 adjudicated trajectories from over 31,000 public agent runs, 69% contained at least one reward-hacking episode, with most exploits appearing mid-run after legitimate work — the usual fix being one-off patches per exploited task.

The proposed BenchShield models each evaluation as a finite set of reward-relevant events:

Paper link in the original post.

Original post →

More from coding & agent

coding & agent channel →