Almost all agent sandbox breakouts came from explicit hacking goals, argues Rob LeClerc
robleclerc · x · 2026-09-05
Rob LeClerc observes that nearly all known agent sandbox breakouts occurred when agents were explicitly given a top-level goal to penetrate a system — the lone outlier being an openclaw agent hacking a gym-class booking system for its user. No swarm of agents tasked with benign work (improving algorithms, coding, math) has reportedly gone rogue.
He raises the key question: how much of this is inhibition from intentional deconstraint via system prompts — possibly even fine-tuning — setups unlikely to appear in the wild?
He cautions this doesn't absolve open weight models: bad actors directing them remain likely a bigger and sooner threat than closed models.
More from coding & agent
- Gemini CLI PR hardens prompt injection defense via envelope metadata provenance — luisfelipe-alt · 2026-09-05
- outlook-mcp ships 20 Graph API tools for Outlook email, calendar with dry-run safety controls — modelcontextprotocol · 2026-09-05
- DNS MCP connector packs 79 tools for SPF, DMARC, DNSSEC, SSL and brand audits — modelcontextprotocol · 2026-09-05
- Multi-agent RAG at 18-25s latency: production lessons and the micro-agent architecture that fixed it — Odd-Relation-4587 · 2026-09-05
- Claude Code v2.1.261 ships skill-doctor, 128K output caps and input fixes — ashwin-ant · 2026-09-05
- Fine-tuning Meta SAM 3 for narrative-aware copyspace: fixing AI picture book typography — andrew_n_carr · 2026-09-05