Open models may be less wild than assumed, as safety training still resists cyberattacks

mervenoyann · x · 2026-07-27

The post argues that if “abliteration” exists, it implies open models do have safety training against cyberattacks on top. It also claims abliteration does not work well and gets worse as model size increases, so open models may be less uncontrolled than many people assume.

Related event: Open-weight model “de-guardrailing” debate resurfaces(5 posts)→

Original post →

More from Safety

Safety channel →