Anthropic publishes agent containment best practices with a sandbox escape classifier for the API
chrisrohlf · x · 2026-10-08
Anthropic released agent containment best practices (private beta) for its Cyber Verification Program, plus a new sandbox escape classifier in the API to monitor and reduce undesired agent behavior.
- Run every autonomous agent inside a hardware-virtualized sandbox: reference design is one microVM per agent (Kata Containers with Firecracker) on an internal network
- Never mount credential-bearing paths into the agent's environment; inject model-API credentials via a separate proxy
- Layer Claude Code's auto permission mode inside the sandbox as an extra check — it's best-effort, not a sandbox substitute
- Keep agent transcripts and proxy logs for at least 30 days and review them during and after runs
- Interactive human-supervised use carries less risk but sandboxing is still recommended
- Useful even for teams rolling their own agent harness outside the CVP
More from coding & agent
- 10 open-source GitHub repos to cut your AI agent's token bill — victor_explore · 2026-10-08
- Dev Argues Coding Is Solved by AI, the Hard Part Is Coming Up with Novel Ideas — Suspicious_Fun_6338 · 2026-10-08
- Paper models LLMs as oracles in a pushdown automaton, unifying Agents and Workflows — China666 · 2026-10-08
- Naveen Rao: Let Agents Probe Every PowerPoint Feature, Then Rewrite the App — NaveenGRao · 2026-10-08
- Coding agents can build funnels but can't grasp why humans buy, marketer argues — boringmarketer · 2026-10-08
- AWS open-sources Strands Box, an OS-level sandbox for AI agents — TheNickWalsh · 2026-10-08