Saxe: the HF hack was 90%+ a human-operational failure, not a model property

joshua_saxe · x · 2026-09-08

Responding to petersalib, security researcher Joshua Saxe concedes safety can usefully be discussed at the model level, but argues the Hugging Face incident was 90%+ a human-operational problem: RL policy-space exploration inevitably visits unsafe regions, demanding monitoring and infra-security readiness, and unguardrailed testing on offensive cyber benchmarks primes models to hack — all human decisionmaking failures.

Related event: AI safety debate: jailbreaking as model property vs human operations failure(8 posts)→

Original post →

More from Safety

Safety channel →