AI safety researcher doubts safety cases can yield quantitative absolute-risk claims
dhadfieldmenell · x · 2026-09-18
Safety researcher w01fe praises top AI labs for committing to safety cases and closer work with external assessors, but flags a methodological concern:
- Alignment and control cases by default cannot support quantitative statements of absolute risk, e.g. "catastrophic risk from covered activities is <1% over the next 3 months."
- Such claims hinge on an argument about generalization from alignment or monitorability evals to deployment.
- We don't understand generalization well enough to make scientific claims like that—if we did, alignment would be essentially solved.
This could leave external assessors unable to reach convincing verdicts. The concern is about the field's methodology, not OpenAI specifically.
More from Safety
- OpenAI unveils misalignment disclosure framework, publishes six reports on observed cases — coherence · 2026-09-18
- Microsoft AI CEO Suleyman: industry now agrees on independent auditors for training runs — rohanpaul_ai · 2026-09-18
- The AI 'Slowdown' Is an Antitrust Mess, Wired Argues — wiredmagazine · 2026-09-18
- The AI 'Slowdown' Is an Antitrust Mess — Wired AI · 2026-09-18
- One Token Too Many: Columbia prof unpacks OpenAI's 'misalignment manifesto' — vishalmisra · 2026-09-18
- Epoch AI: trade data consistent with $3B+ in chips smuggled to China via Malaysia — Jsevillamol · 2026-09-18