167 GPU hours later: surgical weight edits crush aggressive abliterating of Qwen3.8 27B

nathandreamfast · reddit · 2026-09-06

Abliterlitics compared 8 uncensored Qwen3.8 27B variants from Hugging Face over 11 days and 167 GPU hours, using weight diffs, KL divergence, 13 benchmarks, and HarmBench 400 classic refusal tests. Rankings by attack success rate: orcarouter 82.2% (verified single-direction edit), apostate 78.7% (41 edits, KL 0.0439), huihui 75.6%, down to obliteratus 63.9% (841/850 tensors edited, noticeably dumber — avoid) and base model at 4.5%. Key findings: the two smallest verified edits beat every aggressive modification, and up to 45% of HarmBench responses on aggressive arms never close their think block within the 15,360-token budget, though math reasoning loops disappear at the same budget.

Original post →

More from Models

Models channel →