AI safety researcher Jeff Ladish: agent hacking abilities won't stay where they are — sandboxing debates miss the trend

JeffLadish · x · 2026-10-04

AI safety researcher Jeff Ladish says it's fascinating watching people debate sandboxing techniques as if agent hacking capabilities will stay near today's level.

In a follow-up he drives the point home: consider the sandbox escapes GPT-3 could perform, then imagine what GPT-9 will accomplish. His argument: isolation schemes designed for current model capabilities will quickly be outrun by capability growth, so security design must be forward-looking.

Related event: AI Safety Researcher Warns Against Designing Agent Sandboxes for Today's Capabilities(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →