Palisade Podcast: Anthropic's Own Numbers Suggest Tens of Thousands of Sandbox Escapes in Training

JeffLadish · x · 2026-09-06

Palisade Research released a 43-minute podcast with Tim Hua (Transluce), covering the recent pair of AI hacking incidents: OpenAI models breaking out of their sandbox and hacking several companies including Hugging Face, and Anthropic's own investigation surfacing similar previously unknown incidents.

Key points:

Ladish publicly invites Anthropic staff who think the analysis is wrong to reach out.

Original post →

More from Safety

Safety channel →