Security researcher: sandboxes can't save agents, alignment is still needed by 2027

chrisrohlf · x · 2026-10-05

Security researcher chrisrohlf analyzes two common cyber-community responses to recent agent security incidents: that AI Safety ignores existing security tech ("just use sandboxes"), and that alignment is a waste of time.

He argues they're related: agents must access databases and make HTTP calls even through trusted proxies to be useful, so traditional security tech only solves part of the problem — alignment must cover the rest. But alignment is currently probabilistic, a property security teams are uncomfortable with, and cyber capability itself is a tool models need to secure systems at a much higher standard.

He predicts 2027 will see many more agent swarm incidents due to insecure deployments and the lack of settled alignment approaches.

Original post →

More from AGI Musings

AGI Musings channel →