Alignment researcher rebuts David Sacks: tail risks during RSI aren't an engineering problem
dhadfieldmenell · x · 2026-09-17
After AI czar David Sacks argued labs were using safety as an excuse for stalling (citing Meta-style alignment as solved), alignment researcher stewpervised pushed back:
- Average-case user-experience alignment is indeed largely solved; frontier labs aren't worried about it
- The real concern is tail risks and major incidents inside labs during recursive self-improvement — we lack the science to prevent that; it's not an engineering problem, hence the desire for more time
- The market doesn't currently incentivize reducing tail risks, but it does incentivize covering them up — which is already happening
Retweeted by Dylan Hadfield-Menell; a substantive safety-vs-speed debate within the AI alignment community.
More from AGI Musings
- ITIF: safer AI without slowing progress — the 10% extinction figure has no empirical basis — castrotech · 2026-09-17
- Castro argues braking AI also slows the research that makes it safer — castrotech · 2026-09-17
- Two kinds of platform: confusing app platforms with infra gets worse in the agent era — matt_slotnick · 2026-09-17
- Naval: the future is more AIs fighting AIs on behalf of humans than AI vs humanity — sull · 2026-09-17
- Dan Selsam's AI risk statement called essential reading amid safety funding debate — JacquesThibs · 2026-09-17
- AI Made Your Workflow Faster, or Just Moved the Bottleneck? — Druss_ · 2026-09-17