Open models may be less wild than assumed, as safety training still resists cyberattacks
mervenoyann · x · 2026-07-27
The post argues that if “abliteration” exists, it implies open models do have safety training against cyberattacks on top. It also claims abliteration does not work well and gets worse as model size increases, so open models may be less uncontrolled than many people assume.
Related event: Open-weight model “de-guardrailing” debate resurfaces(5 posts)→
More from Safety
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23
- Open-source advocates call doom narratives a regulatory moat against open weights — AlexTensor · 2026-09-23
- AI safety will follow engineering tradition: formal proofs for simple cases, evals for complex — burny_tech · 2026-09-23
- Stochastic Parrots authors rebut AI-pause letter: focus on present harms, not sci-fi risk — marigo · 2026-09-23
- Devs mock labs' cyber-enabled Claude/GPT testing as 'felonies sold as safety research' — ctjlewis · 2026-09-23
- Okta launches Human Principal, binding AI agents to verified humans via World ID — BecauseCulture · 2026-09-23