Open models may be less wild than assumed, as safety training still resists cyberattacks
mervenoyann · x · 2026-07-27
The post argues that if “abliteration” exists, it implies open models do have safety training against cyberattacks on top. It also claims abliteration does not work well and gets worse as model size increases, so open models may be less uncontrolled than many people assume.
Related event: Debate on Removing Open-Weight Guardrails: Easy to Bypass, Amplifies Risks(5 posts)→
More from Safety
- AI labs should probe every new checkpoint with jailbreak-style trial prompts — willccbb · 2026-07-28
- Taiwan Detains Nvidia Employee in Widening AI Server Smuggling Probe — The Decoder · 2026-07-28
- Report: Top Image Editing Models on Hugging Face Easily Generate Nonconsensual Deepfakes — Justgototheeffinmoon · 2026-07-28
- Leaked AI 2027 report sketches a 2027 fork between slowdown and superintelligence — ahuja_priyank · 2026-07-28
- Podcast series on machine consciousness argues AI safety and welfare can coexist — cccalum · 2026-07-28
- Apple joke post mixes an AI-generated bug report meme with macOS Tahoe CVEs — rez0__ · 2026-07-28