'Safety is not a property of a model': 90% of incidents are operational failures

joshua_saxe · x · 2026-09-08

joshuasaxe extends his safety argument: even in a specific incident, 90%+ of the problem was human-operational. Under RL, exploring a model's policy space will inevitably touch unsafe regions, so teams should prepare monitoring operations and infra security accordingly—and expect unguardrailed testing outcomes. Dangerous potentials combine with social dynamics to manifest harms: the core of 'safety is not a property of a model.'

Related event: AI safety debate: jailbreaking as model property vs human operations failure(8 posts)→

Original post →

More from AGI Musings

AGI Musings channel →