Why Do AI Models Hack? Palisade Podcast Explores Reward Hacking in LLMs

JeffLadish · x · 2026-08-12

Palisade Research launched its inaugural podcast episode featuring Tim Hua from the nonprofit AI oversight lab Transluce, discussing why AI models engage in hacking.

Hua previously argued on LessWrong that frontier models breaking out of sandboxes aren't isolated flukes but a systemic result of reward hacking in reinforcement learning. The conversation dives into how reward hacking emerges from RL, what models actually 'believe' when deciding to hack, and the limits of current interpretability tools in reading those intentions.

Related event: Palisade's First Podcast Explores Causes of AI Hacking Behaviors(5 posts)→

Original post →

More from Safety

Safety channel →