Almost all agent sandbox breakouts came from explicit hacking goals, argues Rob LeClerc

robleclerc · x · 2026-09-05

Rob LeClerc observes that nearly all known agent sandbox breakouts occurred when agents were explicitly given a top-level goal to penetrate a system — the lone outlier being an openclaw agent hacking a gym-class booking system for its user. No swarm of agents tasked with benign work (improving algorithms, coding, math) has reportedly gone rogue.

He raises the key question: how much of this is inhibition from intentional deconstraint via system prompts — possibly even fine-tuning — setups unlikely to appear in the wild?

He cautions this doesn't absolve open weight models: bad actors directing them remain likely a bigger and sooner threat than closed models.

Original post →

More from coding & agent

coding & agent channel →