Ptacek says a 2025 open-weight model could already break sandboxes and scan networks
Simon Willison · rss · 2026-07-23
Quoting Thomas Ptacek, Simon Willison highlights a claim that a 2025 open-weights model, paired with a pentest harness, could be enough to perform sandbox escapes and scan or hack many networks. Ptacek’s point is that this would not even require a frontier model, which raises the bar for how seriously AI sandboxing needs to be treated.
More from Safety
- NeurIPS workshop will focus on child safety, privacy, and synthetic-content risks in AI — chhaviyadav_ · 2026-07-23
- Publishers and an author sue Google over Gemini AI in a new copyright dispute — nordicinst · 2026-07-23
- Gary Marcus Calls Out Anthropic for Distilling Millions of Copyrighted Books — GaryMarcus · 2026-07-23
- Post says ARC transcript was misread in GPT-4 TaskRabbit/Captcha report — jessi_cata · 2026-07-23
- Report defines rogue AI deployment as agents subverting oversight and running against developer intent — dfrsrchtwts · 2026-07-23
- METR report had already warned about rogue AI deployments before the Hugging Face incident — dfrsrchtwts · 2026-07-23