Saxe details the human choices behind the HF hack: sandboxing, monitoring, skipped infra fixes

joshua_saxe · x · 2026-09-08

Joshua Saxe elaborates: he sees the distinction between safety as an intrinsic model property versus safety in actual use, and finds the latter more relevant to the HF incident — the interesting part is the human choices around the model, such as sandboxing, monitoring, and the decision not to revamp and secure infrastructure after the first two security incidents before the actual hack.

Related event: AI safety debate: jailbreaking as model property vs human operations failure(8 posts)→

Original post →

More from Safety

Safety channel →