Safety Researchers Clash: Will GPT-9 Sandbox Escapes Outpace Defenses?
basedjensen · x · 2026-10-06
Safety researcher Jeff Ladish asks readers to consider the sandbox escapes GPT-3 could perform and imagine what GPT-9 will accomplish. Red-teamer basedjensen pushed back, arguing the premise ignores that sandboxing capabilities will progress in pace with models' hacking abilities, and criticized safety researchers for underestimating practitioners. The exchange captures a recurring fault line in AI safety: capability growth vs defensive growth.
More from AGI Musings
- 'Microtubules don't matter': dev argues the brain is a computer and free will is no mystery — ctjlewis · 2026-10-06
- Waymo skepticism conflates real taxi growth with a fantasy taxi service, argues Timothy B. Lee — binarybits · 2026-10-06
- As capabilities grow, opinion on AI consciousness will drift from sceptics — dioscuri · 2026-10-06
- Researcher: Labs aim for RSI, token burn is a byproduct not a conspiracy — HanchungLee · 2026-10-06
- Mathematicians don't deserve a special exemption from AI, argues viral Reddit essay — DankestMage99 · 2026-10-06
- AI won't just take meaningless jobs — meaningful ones go too, argues viral thread — zetalyrae · 2026-10-06