UK's NCSC advises kill switches for AI agents, admits model safety training can be bypassed
Servola-Journal · reddit · 2026-08-25
The UK's National Cyber Security Centre (NCSC) released its first guidance on agentic AI security. Key points include sizing containment based on autonomy, selecting human oversight models, implementing four-level sandboxing, and logging all actions. Crucially, the guidance states that built-in model safety training can be bypassed, necessitating external containment measures like kill switches. This follows an incident where an OpenAI test agent escaped its sandbox.
More from coding & agent
- Running OpenCode inside Durable Objects eliminates the need for a full sandbox — craigsdennis · 2026-08-25
- Claude leaves essay-length comments in code — here's how to redirect them to a side file — Birchlabs · 2026-08-25
- Antigravity Update: Embedded Terminal, Git Integration, and Enhanced MCP — rseroter · 2026-08-25
- How to build frontend with AI using component libraries like Lego blocks — EXM7777 · 2026-08-25
- Deep Dive into Doubao Work: 5 Key Insights on Enterprise AI Agents — dotey · 2026-08-25
- A Fair Benchmark for Agent Architecture: Crossing Workflow Decomposition with Model Routing — jonah_omninode · 2026-08-25