Safety researcher: AI companies' safety cases lack detail; here's what to disclose
thlarsen · x · 2026-09-13
- Responding to the proposal that frontier model developers create and stick to safety cases verified by third parties, safety researcher thlarsen calls it a major step forward but says companies' safety cases lack the detail to be trusted, and the visible analysis is dubious — evidence doesn't support top-level claims that risk is low.
- His disclosure wishlist includes: the model's opaque serial depth; scope of control/monitoring measures across evals, training and internal use (noting OpenAI apparently did not monitor evals during the HF incident); a complete list of incidents monitoring has caught and ones it missed; and current beliefs about AI model motivations and their basis.
More from Safety
- SALT, MAD and containment did slow nuclear proliferation — a case for the AI containment analogy — aronchick · 2026-09-13
- The Hard Part of Slowing AI Is Verifying Everyone Else Is Slowing Down — SiteSpecialist6295 · 2026-09-13
- Verdon predicts regulators' next move: making open-source models illegal — beffjezos · 2026-09-13
- Extropic CEO warns regulation is strangling frontier AI and open source — beffjezos · 2026-09-13
- CAIS: Securing AI Weights From Nonstate Actors Will Take 6-12 Months — naturecomputes · 2026-09-13
- CAIS answers 'what to do during an AI slowdown': containment, propensities, robustness, institutions — naturecomputes · 2026-09-13