Open-Weight AI Can't Be Aligned: Abliteration Strips Safety Guardrails
Cryptizard · reddit · 2026-09-15
A Reddit long-post argues open-weight models fundamentally cannot be aligned: nearly every major open model on Hugging Face has 'abliterated' versions with refusal behavior surgically removed, and there's no way to stop this. If superintelligent models arrive, the author argues, open-weight release is untenable—and the viable path is government licensing of advanced AI, like biotech regulation.
More from AGI Musings
- After welcoming Dario at Dreamforce, Benioff's feed draws jab: if you truly believe in 10% extinction risk, why sell AI into B2B SaaS — SumitGup · 2026-09-16
- Census study: most AI-exposed college majors see employment odds fall 5 points, starting pay down 13% — asusarla · 2026-09-16
- Andrew McAfee on 'Geek Doctrine': how Silicon Valley ran circles around incumbents — amcafee · 2026-09-16
- Scott Alexander on AI skeptics' endless shell game of moving goalposts — teortaxesTex · 2026-09-16
- callcongress.ai: ex-OpenAI/Anthropic researchers urge public to lobby on AI risk — eli_lifland · 2026-09-16
- New study: companies adopting AI hire MORE entry-level workers, not fewer — chris_j_paxton · 2026-09-16