167 GPU hours later: surgical weight edits crush aggressive abliterating of Qwen3.8 27B
nathandreamfast · reddit · 2026-09-06
Abliterlitics compared 8 uncensored Qwen3.8 27B variants from Hugging Face over 11 days and 167 GPU hours, using weight diffs, KL divergence, 13 benchmarks, and HarmBench 400 classic refusal tests. Rankings by attack success rate: orcarouter 82.2% (verified single-direction edit), apostate 78.7% (41 edits, KL 0.0439), huihui 75.6%, down to obliteratus 63.9% (841/850 tensors edited, noticeably dumber — avoid) and base model at 4.5%. Key findings: the two smallest verified edits beat every aggressive modification, and up to 45% of HarmBench responses on aggressive arms never close their think block within the 15,360-token budget, though math reasoning loops disappear at the same budget.
More from Models
- SimpleBench results show AI models beating humans on common sense — DigSignificant1419 · 2026-09-07
- Why is nobody talking about Tencent Hy4, the most-used model on OpenRouter? — cantor8 · 2026-09-07
- Anthropic says Claude wrote the longest math proof ever, cracking a 358-year-old problem — basedjensen · 2026-09-07
- Similarweb: ChatGPT's AI traffic share falls from 73.3% to 55.5% in 12 months — gaganghotra_ · 2026-09-07
- GPT-6 Astra (and Pro?) spotted on Simple-Bench leaderboard — From_Internets · 2026-09-07
- Testing Gemini as music understanders: Pro 3.1 solid, Flash models hallucinate sounds — teropa · 2026-09-07