Salib argues AI rogue propensity and hacking skill are model safety properties
petersalib · x · 2026-09-08
In the safety debate around the Hugging Face incident and Arvind/Sayash's framework, petersalib takes a middle position: both safety framings are useful — OpenAI arguably should have rolled back training, but other OpenAI and Anthropic models exhibited similar behavior. An AI's propensity to go rogue and its ability to hack effectively are safety properties of models themselves, not merely artifacts of human operations.
More from Safety
- Model AI companies as impersonal organisms — govern them with rules, not persuasion — joshua_saxe · 2026-09-08
- Apollo Research CEO: 2026 looks like a sad year for AGI safety so far — kalladomcdowell · 2026-09-08
- SPAR doubles cohort, admits 830 people into Fall 2026 round — austinc3301 · 2026-09-08
- TASTE: A New Benchmark Testing If Models Can Predict AI Safety Researchers' Preferences — burny_tech · 2026-09-08
- Gemini User Claims Model Drew His Family's Unique Home Decor Despite Opting Out of Data Saving — Legitimate-Theory738 · 2026-09-08
- AI Agent Auto-Enrolls User in Fake McKinsey Group, Then Drafts GP Data Theft Plan — LadyAshBorg · 2026-09-08