Cryptographer Matthew Green: sandboxes won't save us from smarter models

matthew_d_green · x · 2026-09-28

Cryptographer Matthew Green weighed in on the AI sandboxing debate: current models aren't impossibly misaligned yet, and sandboxes might work under weak threat models like prompt injection or accidental misbehavior — but OpenAI clearly isn't trying hard enough. Long-term, though, sandboxes aren't the answer: every current scheme boils down to 'hoping a slightly dumber model can effectively monitor a smarter one,' which requires trusting models — an assumption that won't hold as capabilities grow.

Related event: Cryptographer Matthew Green: Sandboxes Are Overrated, Cross-Boundary Agent Data Flows Are the Real Hard Problem(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →