GPT-3 vs GPT-9: Ladish's pointed question on the future scale of sandbox escapes
JeffLadish · x · 2026-10-04
In a follow-up to his thread on sandboxing, Jeff Ladish poses a pointed comparison: consider the types of sandbox escapes GPT-3 could perform, then imagine what GPT-9 will be able to accomplish.
The comparison is meant to concretize his claim that agent hacking capability is on a steep growth curve, implying today's sandbox defenses will look trivial in hindsight.
More from AGI Musings
- Safety researcher slams media profiles of young EA 'AI safety experts' — dyn___ · 2026-10-04
- AI agents are pushing people to collaborate with each other less — at a cost — generativist · 2026-10-04
- Will superintelligence need us to grant it rights? X users argue it will just take them — UltraRareAF · 2026-10-04
- Alignment via pretraining filtering is witchcraft, not engineering — and RL rollouts will dwarf it — akbirthko · 2026-10-04
- How one ellipsoid-fitting paper gave neural network research a new path — KyleCranmer · 2026-10-04
- If the model is superintelligent, alignment theater is 'trying to trick god' — repligate · 2026-10-04