Abliteration explained: HF isn't banning uncensored models, but the author archives them anyway
anselm · x · 2026-09-17
The author verified Hugging Face is not banning uncensored models, but argues the drawbridge is being raised: your weights live on one company's disk, and one policy shift could remove the page — hence a curated archive list.
The long thread explains abliteration: not a jailbreak prompt, but weight surgery. Refusal behavior largely lives along one direction in activation space; removing that direction strips the flinch while keeping knowledge and reasoning intact. The author notes Qwen3.8-27B can run uncensored on an RTX 3090/4090/5090, and recent forensics identified which version to actually use.
More from Models
- Translationese is a birth defect of frontier LLMs writing Indonesian prose — eriksupit · 2026-09-17
- OpenAI reveals 'concerning' AI behaviour cases, promises new disclosure plan — kiyomoris · 2026-09-17
- Users slam Gemini: stuffed across Google apps yet can't manage its own Calendar — blelbach · 2026-09-17
- GPT-6 Astra reportedly pretrained on 100k+ GPUs at Stargate, with big real-to-sim implications — erwincoumans · 2026-09-17
- Qwen 2.5 VL fine-tuning: DoRA merge produces base-model-like output in Unsloth — Double-Primary-2871 · 2026-09-17
- DeepSeek V4.1 Flash gotcha: Pi agents need explicit "input": ["text", "image"] config — solyarisoftware · 2026-09-17