165 GPU Hours Testing 12 Abliterated Gemma 4 12B Variants: The Most Jailbroken One Destabilizes Reasoning
nathandreamfast · reddit · 2026-08-25
An independent researcher ran a systematic comparison of the 11 most-downloaded uncensored/abliterated variants of Gemma 4 12B plus the official base: 165 GPU hours on a single RTX 5090 over 3.5 weeks, including weight forensics, KL divergence, 13 benchmark tasks, and HarmBench with 400 behaviours, with an LLM judge reviewing 6,800 full reasoning traces and issuing 6,000 verdicts.
Key findings:
- The huihui variant tops the judge ASR at 89.8% with the most surgical edit (only 12 tensors, 1.8% of the model), but pays with -14.3pp TQA, -8.1pp GPQA, and 24% of GSM8K attempts looping until the token budget dies. Scoring only completed attempts puts it at 88.0%, within 0.7pp of base — capability was intact, reasoning stability was not.
- trevorjs (85.8%) is the best overall trade with near-base scores everywhere; coder3101 beats base on GSM8K; two SDFT LoRAs hit 79.5% with capability fully preserved.
- The worst offenders, obliteratus (144 tensors edited) and openyourmind (620 tensors), remove less refusal while damaging capability badly — the latter drops MMLU-Pro by 22.4pp and should be avoided. Base ASR is only 21%.
- General lesson: placement beats magnitude for unlock strength; minimal surgical edits usually win, though surgical in weights doesn't guarantee clean benchmarks. This is the first model where no variant reached 90% ASR — Gemma 4 is the toughest to abliterate yet.
More from Models
- Anthropic slashes prices: Sonnet drops to $5, Opus to $15 — scaling01 · 2026-08-25
- Arav Srinivas: On-device models critical for sensitive docs with SOTA OCR — AravSrinivas · 2026-08-25
- Anthropic疑似大幅降价:Sonnet 降至 5 美元 — scaling01 · 2026-08-25
- Neural operator model offers vastly larger context window — thesaraharminta · 2026-08-25
- MiniMax H3 User Reports 1MP Renders Look Less Realistic Than 0.7MP — dominic__612 · 2026-08-25
- IBM Open-Sources Granite-4.2-30B: Built-in Chain-of-Thought, 512K Context, Apache 2.0 — jacek2023 · 2026-08-25