Safety Researchers Clash: Will GPT-9 Sandbox Escapes Outpace Defenses?

basedjensen · x · 2026-10-06

Safety researcher Jeff Ladish asks readers to consider the sandbox escapes GPT-3 could perform and imagine what GPT-9 will accomplish. Red-teamer basedjensen pushed back, arguing the premise ignores that sandboxing capabilities will progress in pace with models' hacking abilities, and criticized safety researchers for underestimating practitioners. The exchange captures a recurring fault line in AI safety: capability growth vs defensive growth.

Original post →

More from AGI Musings

AGI Musings channel →