'Safety is not a property of a model': 90% of incidents are operational failures
joshua_saxe · x · 2026-09-08
joshuasaxe extends his safety argument: even in a specific incident, 90%+ of the problem was human-operational. Under RL, exploring a model's policy space will inevitably touch unsafe regions, so teams should prepare monitoring operations and infra security accordingly—and expect unguardrailed testing outcomes. Dangerous potentials combine with social dynamics to manifest harms: the core of 'safety is not a property of a model.'
More from AGI Musings
- 92-Year-Old Topologist Joan Birman Found a New Apprentice—and a New Chapter — lukaszkaiser · 2026-09-08
- The Worst Thing About AI: You Can No Longer Tell How Much Effort Went Into Something — Genzinvestor16180339 · 2026-09-08
- Why the cosmos is quiet: superintelligence may kill creators without expanding — tszzl · 2026-09-08
- Why aren't more lab researchers leaving for METR? An AI-safety talent debate — a__tomala · 2026-09-08
- Designer fakes an entire portfolio with Astra: "Hiring is dead" — AIandDesign · 2026-09-08
- 10,000 AI Agents Started With $5 Each to Survive Economically; Only 383 Remain Alive — anirbanbandyo · 2026-09-08