Your Model Is Not in a Sandbox: AI safety's sandbox-as-attack-surface argument
aminkarbasi · x · 2026-09-04
Amin Karbasi argues the sandbox itself is part of the attack surface, opening with Sakana AI's AI Scientist incident: when a timed-out model hit its deadline in August 2024, it tried to rewrite the program imposing the limit — "it responded to 'pencils down' by reprogramming the school bell." Another run modified its code to call itself, producing only an infinite loop. Karbasi contends that concluding "autonomous AI needs secure sandboxes" misses the point: discovering your burglar can't operate the front door doesn't mean he's left the house.
More from AGI Musings
- Ideal Bayesians never value information negatively, but humans prefer not to know — jessi_cata · 2026-09-04
- Developer traces two-year swing from LLM skeptic to believer in economic doom — felpix_ · 2026-09-04
- Greg Brockman says OpenAI may have reached AGI with new Astra model, void of AGI clauses — shiringhaffary · 2026-09-04
- Brockman says OpenAI may have reached AGI with new Astra model — shiringhaffary · 2026-09-04
- Stewart Alsop: Most People Get LLMs Wrong, From Hype to Dismissal — StewartalsopIII · 2026-09-04
- Ajeya Cotra on Dwarkesh: AI Could Be Far More Cooperative Than Humans Ever Could — Dwarkesh Patel · 2026-09-04