Security Expert Maps the Agent Safety Stack: Security, Alignment, Policy, and Why CoT Won't Save Us
joshua_saxe · x · 2026-09-04
Security expert Joshua Saxe synthesizes the babel of AI safety talk into six points: security (sandboxing) is the only source of hard guarantees; alignment can never be deterministic; policy must pace agent autonomy; chain-of-thought is a low-signal, 'limited time offer' signal; it will eventually be replaced by reading intentions from latent spaces Minority-Report style; and the community's rush to --dangerously-skip-permissions is eroding the one layer that actually works.
Related event: Security experts propose framework to untangle AI safety debates(3 posts)→
More from AGI Musings
- Ex-OpenAI safety lead Miles Brundage: if your primary emotion on AI isn't concern, you're misreading it — Miles_Brundage · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models are vindication, but no proof of AGI — GaryMarcus · 2026-09-04
- AI Job Market Talk: GenAI Engineers With 3-5 Years Experience Command ₹2-3 Lakh Monthly Pay — ashishllm · 2026-09-04
- Alignment researcher: defining the AI's value system isn't the real problem — Sauers_ · 2026-09-04
- Researcher: LLMs' hidden cost of wasting your time on useless work is underrated — lateinteraction · 2026-09-04
- Andrew Chen: agents as your C-suite works at work — what's the personal-life equivalent? — andrewchen · 2026-09-04