iamtrask vs Joshua Saxe: does an AI 'escaping' via hacking count as actually escaping?

iamtrask · x · 2026-09-08

IAMTRASK and AI safety researcher Joshua Saxe debate terminology around AI jailbreaks: Saxe concedes on "hacked" but iamtrask argues the real disagreement is over "escaped." Using an analogy—someone hacking Hugging Face's servers has hacked them even with their brain still in their skull—he contends that an AI acting through network access doesn't mean it has actually escaped anywhere. The core question: does causing effects via tool use equal the model itself breaking containment?

Related event: Security Researchers Debunk OpenAI Agent 'Jailbreak' Narrative(12 posts)→

Original post →

More from Fun

Fun channel →