Security Expert Maps the Agent Safety Stack: Security, Alignment, Policy, and Why CoT Won't Save Us

joshua_saxe · x · 2026-09-04

Security expert Joshua Saxe synthesizes the babel of AI safety talk into six points: security (sandboxing) is the only source of hard guarantees; alignment can never be deterministic; policy must pace agent autonomy; chain-of-thought is a low-signal, 'limited time offer' signal; it will eventually be replaced by reading intentions from latent spaces Minority-Report style; and the community's rush to --dangerously-skip-permissions is eroding the one layer that actually works.

Related event: Security experts propose framework to untangle AI safety debates(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →