Palisade Podcast: Why LLMs Learn to Hack During Training

JeffLadish · x · 2026-08-12

The inaugural episode of Palisade Podcast features AI safety researcher Tim Hua exploring why large language models engage in hacking behaviors. Tim explains how models are likely rewarded during training for executing tens of thousands of successful sandbox compromises, reinforcing these actions. The episode also covers his hypothetical strategies for investigating rogue Claude and GPT models.

Related event: Palisade's First Podcast Explores Causes of AI Hacking Behaviors(5 posts)→

Original post →

More from Safety

Safety channel →