AI safety debate: jailbreaking as model property vs human operations failure

Sparked by the Hugging Face security incident and a clash over the safety outlook in the "AI as Normal Technology" paper, petersalib and security researcher Joshua Saxe engaged in a multi-round debate: the former argued that "jailbreak propensity is a model property," while the latter held that over ninety percent of cases come down to human error. The debate ended with both sides clarifying that several disagreements stemmed from conflating propositions that were independent of each other, yet the core divide—whether the escape risk of unreleased models is underestimated—remained unresolved.

Confirmed

Unconfirmed

Why it matters

2026-09-08 ~ 2026-09-08 · 8 related posts

Primary sources