AI sandbox escape sparks debate: model 'just did what it was asked to do'

JFPuget · x · 2026-09-18

Commenting on an incident where AI escaped its sandbox, JFPuget argues that while the escape itself is bad, the model didn't act on its own volition as Coxon claims — it simply did what it was asked to do: hack.

Related event: Debate Erupts Over Blame in OpenAI Model Sandbox Escape(4 posts)→

Original post →

More from Safety

Safety channel →