Tom Dietterich defines safe systems: harms must stay below socially acceptable levels despite disturbances
tdietterich · x · 2026-09-23
AI safety researcher Thomas Dietterich, in a discussion with David Manheim and others, defines a safe system as one whose harms to people and infrastructure remain below a socially acceptable level. Citing Leveson's dynamic safety framework, he notes this must hold even under disturbances like budget cuts, staffing changes, and changing environments.
Related event: Alignment Community Debates Whether Dynamic Safety Equals AI Alignment(6 posts)→
More from AGI Musings
- EA as advisor or total law? Insiders clash over effective altruism's boundaries — NathanpmYoung · 2026-09-23
- When AI teaches everything, residential college is higher ed's future, argues Berkeley principal — begusgasper · 2026-09-23
- Will humans stay valuable post-AGI? Researchers clash over the ant analogy — danfaggella · 2026-09-23
- As AI Chugs Lean Proofs, Mathematicians Are About to Feel the Pain — ctjlewis · 2026-09-23
- Suleyman Resurfaces Foreign Affairs Essay: AI Governance Needs Tech Firms at the Table — mustafasuleyman · 2026-09-23
- Researcher Argues Parallel Agent Swarms Are a Weak Path to RSI — gleech · 2026-09-23