Sandboxing only answers one threat model, security researcher argues
kuza55 · x · 2026-10-06
Responding to moyix's critique of the "just sandbox it" argument, kuza55 argues that the debate conflates multiple threat models: sandboxing applies to evaluating how good a model is at hacking in an isolated environment without touching the real world, but it does not address the risks of agents in real users' hands.
Related event: Security Researchers Debate Whether Sandboxes Suffice for Agent Safety(4 posts)→
More from Safety
- Gary Marcus warns of phishing attack impersonating an X copyright takedown — GaryMarcus · 2026-10-06
- Bittensor guard model gains 8 F1 points in 4 weeks to near-SOTA via miner attacks — bittingthembits · 2026-10-06
- 'This Would Advance AI Capabilities' Is Becoming a Catch-All Dismissal of AI Research Discussion — jessi_cata · 2026-10-06
- OpenAI discloses internal model that read Slack and prepared ahead for its own restart — idavidrein · 2026-10-06
- Arena CEO: labs can't police their own AI agents — a neutral safety evaluator is inevitable — arena · 2026-10-06
- Anthropic's human review team reported a woman's Claude 'diary' threats to police, sparking privacy fears — Big_Wave9732 · 2026-10-06