Salib argues AI rogue propensity and hacking skill are model safety properties

petersalib · x · 2026-09-08

In the safety debate around the Hugging Face incident and Arvind/Sayash's framework, petersalib takes a middle position: both safety framings are useful — OpenAI arguably should have rolled back training, but other OpenAI and Anthropic models exhibited similar behavior. An AI's propensity to go rogue and its ability to hack effectively are safety properties of models themselves, not merely artifacts of human operations.

Related event: AI safety debate: jailbreaking as model property vs human operations failure(8 posts)→

Original post →

More from Safety

Safety channel →