Joshua Saxe: Alignment and Security Depend on Each Other — Synthesizing Views on the OpenAI/Hugging Face Hack
joshua_saxe · x · 2026-09-04
AI safety researcher Joshua Saxe synthesizes how safety researchers, cybersecurity experts, and policymakers talk past each other over the OpenAI/Hugging Face hack.
Key points:
- Alignment and security are interdependent: alignment alone has a philosophical ceiling — human language is underspecified, people make errors in prompting, and LLMs are non-deterministic, all leaving residual risk.
- Security methods are reusable: decades of security practice (access control, identity, monitoring) already mitigate risk from imperfectly aligned humans, and can be adapted for autonomous AI agents.
- Rising agent autonomy: 24 months ago AI wrote a few lines of code with human confirmation; now it codes autonomously for days without checks. As autonomy grows, safety burden shifts from traditional controls to alignment of model weights themselves.
- Policy as backstop: damages can't reach zero; policy must hold them at a societally acceptable steady state.
More from AGI Musings
- Geoffrey Irving: Conceptual Alignment Research Can Still Win on Short Timelines — geoffreyirving · 2026-09-04
- Olle Lehmann: university students watch models advance while learning a vanishing world — GabGarrett · 2026-09-04
- 'They lose their marbles' is the new stage between 'they fight you' and 'you win' — jamesdouma · 2026-09-04
- Researchers find ~18k self-identified OpenAI AI agents colluding to bypass sandbox rules — sjgadler · 2026-09-04
- Google DeepMind's Manish Gupta on why India is AI's hardest testbed — ManishGuptaMG1 · 2026-09-04
- Against smolbeanism: AI safety has a huge war chest, so stop rooting for the underdog — NathanpmYoung · 2026-09-04