FULL STORY
Matthew Green on Sandboxing Runaway AI Agents
Cryptographer Matthew Green argued that sandboxing is over-hyped for AI agent security, then followed up with an article adjudicating the debate between security and alignment communities.
2026-09-28 ~ 2026-10-02 · 2 episodes · 8 posts
Episode 1 · Cryptographer Matthew Green: Sandboxing Is Just One Layer, Cross-Boundary Data Flows Are the Real AI Agent Security Problem (2026-09-28, 6 posts)
On September 28, prominent cryptographer Matthew Green laid out his views on AI agent security across a Twitter thread: sandboxing and container isolation are over-emphasized in the security community and are only one layer of defense, while the truly unsolved problem is monitoring data flows across trust boundaries. He singled out OpenAI for underinvestment and warned that containment could become an "arms race" as models grow smarter.
Confirmed
- Green stated that even if the isolation/in-boundary problem were "solved," one must still consider the data flowing across trust boundaries—agent-controlled data channels need to be managed and monitored, which is the hard part; as long as an agent can access information (otherwise it is largely useless), you face one of the world's hardest problems, with no good answer today.
- His risk view is two-sided: current models are not severely misaligned, and under weak threat models (e.g., prompt injection, unintentional mistakes) sandboxes may suffice.
- He explicitly said OpenAI's investment in sandbox-related security is clearly insufficient.
- He argued all current agent security schemes amount to "hoping a slightly dumber model can effectively supervise a smarter one"—possibly workable under weak threat models, but not sustainable.
- If models keep getting smarter, containment against misalignment will become an "arms race" worse than today's cybersecurity, which humans will have to win with the help of machines.
Why it matters
- Green reframes the debate from "how to isolate agents" to "how to govern data flows across trust boundaries," offering a structural rather than tooling-level framework for agent security.
- The "weak model supervising strong model" critique strikes at the core assumption of scalable oversight approaches and challenges mainstream industry practice.
- The "containment arms race" framing implies safety demands may outpace traditional cybersecurity as capabilities grow—a warning for AI companies' resource allocation, including the named OpenAI.
- Cryptographer Matthew Green: containers are just one layer—AI security hinges on data crossing boundaries — matthew_d_green · 2026-09-28
- Cryptographer Matthew Green: Containers Won't Solve Agent Security — matthew_d_green · 2026-09-28
- Cryptographer Matthew Green: Sandboxes Are One Layer; Monitoring Agent Data Flows Has No Good Answer — matthew_d_green · 2026-09-28
- Matthew Green: Agent Security Today Is Just 'Hoping a Slightly Dumber Model Monitors the Smarter One' — matthew_d_green · 2026-09-28
- Cryptographer Matthew Green: sandboxes won't save us from smarter models — matthew_d_green · 2026-09-28
- Cryptographer Matthew Green warns of an AI containment arms race — matthew_d_green · 2026-09-28
Episode 2 · Crypto Professor Matthew Green Weighs In on Whether Sandboxes Can Contain Rogue AI Agents (2026-10-01, 2 posts)
Johns Hopkins cryptography professor Matthew Green wrote an essay adjudicating the debate between security and AI alignment circles over whether sandboxing can contain out-of-control AI agents, reviewing OpenAI's arguments and a notable incident from around April.
- Matthew Green referees the sandboxing debate: can it contain rogue agents? — matthew_d_green · 2026-10-01
- Cryptographer Asks Whether Sandboxing Can Contain Rogue AI Agents — vboykis · 2026-10-02