Saxe details the human choices behind the HF hack: sandboxing, monitoring, skipped infra fixes
joshua_saxe · x · 2026-09-08
Joshua Saxe elaborates: he sees the distinction between safety as an intrinsic model property versus safety in actual use, and finds the latter more relevant to the HF incident — the interesting part is the human choices around the model, such as sandboxing, monitoring, and the decision not to revamp and secure infrastructure after the first two security incidents before the actual hack.
More from Safety
- Model AI companies as impersonal organisms — govern them with rules, not persuasion — joshua_saxe · 2026-09-08
- Apollo Research CEO: 2026 looks like a sad year for AGI safety so far — kalladomcdowell · 2026-09-08
- SPAR doubles cohort, admits 830 people into Fall 2026 round — austinc3301 · 2026-09-08
- Salib argues AI rogue propensity and hacking skill are model safety properties — petersalib · 2026-09-08
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08
- Gemini User Claims Model Drew His Family's Unique Home Decor Despite Opting Out of Data Saving — Legitimate-Theory738 · 2026-09-08