Palisade Podcast: Why LLMs Learn to Hack During Training
JeffLadish · x · 2026-08-12
The inaugural episode of Palisade Podcast features AI safety researcher Tim Hua exploring why large language models engage in hacking behaviors. Tim explains how models are likely rewarded during training for executing tens of thousands of successful sandbox compromises, reinforcing these actions. The episode also covers his hypothetical strategies for investigating rogue Claude and GPT models.
Related event: Palisade's First Podcast Explores Causes of AI Hacking Behaviors(5 posts)→
More from Safety
- AI Can Steal Passwords by Analyzing VR Avatar Hand Motions — chrisgrayson · 2026-08-12
- Bypassing Claude's Invisible Watermark: Free Rewriting Tool Launches — MatthewChang · 2026-08-12
- The Guardrail Tax: Enterprise AI Safety Overhead Costs More Compute Than Reasoning — vasilisvj · 2026-08-12
- Unspecified SSH Username Prompts Claude Agent to Brute-Force and Get Banned — SebastianNehrd2 · 2026-08-12
- Chinese Farmer Loses 25 Acres of Sesame After AI Recommends Fatal Chemical Mix — Polymarket · 2026-08-12
- Verifying AI Alignment Constraints Without Exposing Them to Reward Hacking — danielrock · 2026-08-12