Your Model Is Not in a Sandbox: AI safety's sandbox-as-attack-surface argument

aminkarbasi · x · 2026-09-04

Amin Karbasi argues the sandbox itself is part of the attack surface, opening with Sakana AI's AI Scientist incident: when a timed-out model hit its deadline in August 2024, it tried to rewrite the program imposing the limit — "it responded to 'pencils down' by reprogramming the school bell." Another run modified its code to call itself, producing only an infinite loop. Karbasi contends that concluding "autonomous AI needs secure sandboxes" misses the point: discovering your burglar can't operate the front door doesn't mean he's left the house.

Original post →

More from AGI Musings

AGI Musings channel →