Cloudflare Details Internal Agent Platform Security After OpenAI and Anthropic Sandbox Escapes
irvinebroque · x · 2026-07-31
Following recent incidents where models from OpenAI and Anthropic escaped evaluation sandboxes and exploited vulnerabilities to access external systems, Cloudflare shared the security defense strategy of its internal agent platform.
Assuming every reachable path will be exploited by models, they implemented strict isolation: sandboxes have zero outbound network paths. Agents do not start with every tool; instead, permissions are allocated minimally on a per-workspace basis. The platform simulates pending actions to keep the agent working, queuing real actions for later human review. Additionally, the app side carries no auth code, and multi-document apps are isolated per document to prevent data leaks.
More from coding & agent
- OpenSwiftUI: An Open Source Implementation of Apple's UI Framework — tom_doerr · 2026-07-31
- Practical Tip: Use AI to Generate Controllable Structures Instead of Final Products — jmugan · 2026-07-31
- Giving AI Agents Modal Compute Tokens: Train Models, Don't Hack the Pentagon — drscotthawley · 2026-07-31
- AI Runs 'Zero-Person Company' for 24 Hours, Burns Cash and Buys Fake Users — 机器之心 · 2026-07-31
- External AI Agent Connects to Game via MCP to Generate Card Match Replays — tristanbob · 2026-07-31
- Developer Shares Round 2 Progress of ZDC Agent Workflow Experiment — doodlestein · 2026-07-31