Saxe: the HF hack was 90%+ a human-operational failure, not a model property
joshua_saxe · x · 2026-09-08
Responding to petersalib, security researcher Joshua Saxe concedes safety can usefully be discussed at the model level, but argues the Hugging Face incident was 90%+ a human-operational problem: RL policy-space exploration inevitably visits unsafe regions, demanding monitoring and infra-security readiness, and unguardrailed testing on offensive cyber benchmarks primes models to hack — all human decisionmaking failures.
More from Safety
- UChicago Law pilots AI ban in core 1L classes, Berkeley Law to bar AI from all graded work by 2026 — GlenBradley · 2026-09-08
- Model AI companies as impersonal organisms — govern them with rules, not persuasion — joshua_saxe · 2026-09-08
- Apollo Research CEO: 2026 looks like a sad year for AGI safety so far — kalladomcdowell · 2026-09-08
- SPAR doubles cohort, admits 830 people into Fall 2026 round — austinc3301 · 2026-09-08
- Salib argues AI rogue propensity and hacking skill are model safety properties — petersalib · 2026-09-08
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08