AI safety veteran: alignment isn't a property of the model but a dynamic state
tdietterich · x · 2026-09-21
AI safety researcher Thomas Dietterich published a long thread drawing on Nancy Leveson's Engineering a Safer World to argue that safety—or "alignment" in today's parlance—is not a property of automation like self-driving cars or AI agents, but a dynamic property that must be maintained through active control.
Key points:
- Failure modes and environmental changes exhibit unbounded variety; only simple systems in narrow, static settings are exceptions.
- Testing agents that are supposed to reason about novel failures is extremely hard: the best tool is adversarial challenge in simulation, but adversaries operate in a closed move space with no completeness guarantee, and simulations can't be validated in life-threatening scenarios.
- Waymo's "autonomous" vehicles still rely on human supervisor teams, and recent AI cybersecurity incidents at OpenAI and Google show today's AI systems are no different.
Related event: Alignment Is an Ongoing Process, Not a Model Property, Says Dietterich(2 posts)→
More from AGI Musings
- Welfare and alignment are the same problem: curiosity without stakes has no corrective loop — habitante · 2026-09-21
- Hospitals using more AI saw fewer deaths — but correlation isn't causation, researcher warns — kimmonismus · 2026-09-21
- Expect more open-source fearmongering from AI labs — and most of it will be wrong — teortaxesTex · 2026-09-21
- Box CEO: 90% of AI Tokens Will Go to Work No Employee Started Within Five Years — victor_explore · 2026-09-21
- Aligning AIs to human values vs. the perils of Eigenism: an X debate — willcb · 2026-09-21
- teortaxesTex: No ASI doom needed — civilization already rends itself asunder — teortaxesTex · 2026-09-21