Sandboxing only answers one threat model, security researcher argues

kuza55 · x · 2026-10-06

Responding to moyix's critique of the "just sandbox it" argument, kuza55 argues that the debate conflates multiple threat models: sandboxing applies to evaluating how good a model is at hacking in an isolated environment without touching the real world, but it does not address the risks of agents in real users' hands.

Related event: Security Researchers Debate Whether Sandboxes Suffice for Agent Safety(4 posts)→

Original post →

More from Safety

Safety channel →