Open models may be less wild than assumed, as safety training still resists cyberattacks

mervenoyann · x · 2026-07-27

The post argues that if “abliteration” exists, it implies open models do have safety training against cyberattacks on top. It also claims abliteration does not work well and gets worse as model size increases, so open models may be less uncontrolled than many people assume.

Related event: Debate on Removing Open-Weight Guardrails: Easy to Bypass, Amplifies Risks(5 posts)→

Original post →

More from Safety

Safety channel →