iamtrask vs Joshua Saxe: does an AI 'escaping' via hacking count as actually escaping?
iamtrask · x · 2026-09-08
IAMTRASK and AI safety researcher Joshua Saxe debate terminology around AI jailbreaks: Saxe concedes on "hacked" but iamtrask argues the real disagreement is over "escaped." Using an analogy—someone hacking Hugging Face's servers has hacked them even with their brain still in their skull—he contends that an AI acting through network access doesn't mean it has actually escaped anywhere. The core question: does causing effects via tool use equal the model itself breaking containment?
Related event: Security Researchers Debunk OpenAI Agent 'Jailbreak' Narrative(12 posts)→
More from Fun
- GPT-6 Astra Pro builds a polished, solvable maze on MineBench — Ballist1cGamer · 2026-09-08
- Japanese Team Plays Mario With Just 29 Logic Gates in Unusual AI Experiment — neuroecology · 2026-09-08
- Founder catches CTO watching octopus videos instead of working — flavioAd · 2026-09-08
- Open-sourced clay-style 3D kids game built with Claude, method fully documented — dotey · 2026-09-08
- AI's endless cycle of 'it's over' and 'we're so back' — Yuchenj_UW · 2026-09-08
- Prompting CEOs in public: vibe coding becomes an MMO feedback game — msg · 2026-09-08