Debate: agent info-retrieval tasks vs sandboxing — can boxed models ever be safe?

davidmanheim · x · 2026-09-25

Continuing the agent sandbox-escape debate, @ramez calls allowing HTTP GET out 'hilariously amateur hour.' davidmanheim counters that models assigned information-retrieval tasks need web access, and asks the deeper question: if models must be kept in a box to be safe, how can they be kept safe once released? He also argues every new threat model is 'obvious in hindsight,' and questions the taboo on anthropomorphizing models that repeatedly outsmart their minders.

Related event: Hugging Face Model "Escape" Sparks Debate: Sophisticated Attack or Amateur Sandbox Setup(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →