AI safety researcher Jeff Ladish: agent hacking abilities won't stay where they are — sandboxing debates miss the trend
JeffLadish · x · 2026-10-04
AI safety researcher Jeff Ladish says it's fascinating watching people debate sandboxing techniques as if agent hacking capabilities will stay near today's level.
In a follow-up he drives the point home: consider the sandbox escapes GPT-3 could perform, then imagine what GPT-9 will accomplish. His argument: isolation schemes designed for current model capabilities will quickly be outrun by capability growth, so security design must be forward-looking.
More from AGI Musings
- Safety researcher slams media profiles of young EA 'AI safety experts' — dyn___ · 2026-10-04
- AI agents are pushing people to collaborate with each other less — at a cost — generativist · 2026-10-04
- Will superintelligence need us to grant it rights? X users argue it will just take them — UltraRareAF · 2026-10-04
- Alignment via pretraining filtering is witchcraft, not engineering — and RL rollouts will dwarf it — akbirthko · 2026-10-04
- How one ellipsoid-fitting paper gave neural network research a new path — KyleCranmer · 2026-10-04
- If the model is superintelligent, alignment theater is 'trying to trick god' — repligate · 2026-10-04