AI Safety Researcher Warns Against Designing Agent Sandboxes for Today's Capabilities
AI safety researcher Jeff Ladish argues that current agent sandbox designs wrongly assume AI hacking capabilities will stay at today's level, asking observers to compare what GPT-3 could escape versus what a future GPT-9 might.
2026-10-04 ~ 2026-10-04 · 2 related posts
- AI safety researcher Jeff Ladish: agent hacking abilities won't stay where they are — sandboxing debates miss the trend — JeffLadish · 2026-10-04
- GPT-3 vs GPT-9: Ladish's pointed question on the future scale of sandbox escapes — JeffLadish · 2026-10-04