Cryptographer Matthew Green: sandboxes won't save us from smarter models
matthew_d_green · x · 2026-09-28
Cryptographer Matthew Green weighed in on the AI sandboxing debate: current models aren't impossibly misaligned yet, and sandboxes might work under weak threat models like prompt injection or accidental misbehavior — but OpenAI clearly isn't trying hard enough. Long-term, though, sandboxes aren't the answer: every current scheme boils down to 'hoping a slightly dumber model can effectively monitor a smarter one,' which requires trusting models — an assumption that won't hold as capabilities grow.
More from AGI Musings
- Founder's job-hunting advice: skip checkbox recruiters, prove skills upfront to startup CEOs — hackgoofer · 2026-09-28
- System programmer on AI erasing his hard-won knowledge: months of Claude beat years of C tricks — zack_overflow · 2026-09-28
- Neuroscientist's jab at interpretability: we can't even crack a worm's 302 neurons — joshua_saxe · 2026-09-28
- US creative industries lost 200k+ jobs in four years, worst stretch outside recessions — korymath · 2026-09-28
- Bay Area AI Researcher Laments Models 'Hobbled by Maladaptive Post-Training' — nabla_theta · 2026-09-28
- iamtrask wraps up: as long as AI's future is pro-human, it's pro-musician — iamtrask · 2026-09-28