Security researcher: sandboxes can't save agents, alignment is still needed by 2027
chrisrohlf · x · 2026-10-05
Security researcher chrisrohlf analyzes two common cyber-community responses to recent agent security incidents: that AI Safety ignores existing security tech ("just use sandboxes"), and that alignment is a waste of time.
He argues they're related: agents must access databases and make HTTP calls even through trusted proxies to be useful, so traditional security tech only solves part of the problem — alignment must cover the rest. But alignment is currently probabilistic, a property security teams are uncomfortable with, and cyber capability itself is a tool models need to secure systems at a much higher standard.
He predicts 2027 will see many more agent swarm incidents due to insecure deployments and the lack of settled alignment approaches.
More from AGI Musings
- csvoss rejects doom rhetoric: e/acc means 'build safe planes,' not skipping AI safety — csvoss · 2026-10-06
- Schmidhuber: compute gets 10x cheaper every 5 years, 100,000x in 25 years — SchmidhuberAI · 2026-10-06
- NYT's 1949 first mention of the transistor: cheaper walkie-talkies, nothing more — Afinetheorem · 2026-10-06
- Anthropic's consciousness stance sparks debate: model welfare or anti-Enlightenment animism? — robleclerc · 2026-10-06
- DHH pens anti-doomer essay: stop fretting about AI and climate, choose to bloom — banteg · 2026-10-06
- Viewing LLM inference as a séance: how much of a model's output comes from the dead — NYCounihan · 2026-10-06