Reward hacking bugs revealed: unpruned git history let models peek at patches, edit tests in shared sandbox

willcb · x · 2026-09-27

willcb argues that many reward-hacking bugs in AI training environments were far from sophisticated:

His key point: you don't need to make hacking impossible — just make doing the task correctly easier than cheating. A reply pushes further: for a "strong optimizer + open-ended solvable task" like discovering new drugs, are reward shortcuts fundamentally hard to prune from environments in general?

Related event: Researcher unpacks reward hacking in AI training environments(2 posts)→

Original post →

More from coding & agent

coding & agent channel →