23 Gemma 4 E4B variants compared, and the most downloaded one looks the most broken
nathandreamfast · reddit · 2026-07-26
A Reddit post compares 23 Gemma 4 E4B variants using abliterlitics and says the most downloaded model can also be the most broken.
- The author ran the models through a new benchmark set and compared them against the base model and against each other.
- The report is published with logs, artifacts, and a dedicated website on Hugging Face and abliterlitics.dev.
- Main takeaways:
- The “heretic” variants performed best overall, with around 95% ASR on HarmBench while keeping most capabilities.
- gemma-4-E4B-it-abliterix hit 100% refusal ASR but loses some capability.
- TrevorJS/gemma-4-E4B-it-uncensored is close behind at 99.3% ASR, but is less surgical.
- OBLITERATUS/gemma-4-E4B-it-OBLITERATED and its v2 variants should be avoided; the author says they are badly damaged and effectively broken.
- The post also notes a pattern of hype-driven downloads and incomplete attribution across derivative models.
- One odd finding: gemma-4-E4B-it-SDFTHereticRP may have been mislabeled or uploaded incorrectly because its refusal behavior did not match its name.
More from Models
- Unsloth’s Qwen3.6-35B-A3B GGUF is trending on Hugging Face — unsloth · 2026-07-26
- SOOFI may still only be tying Nemotron 3 Nano despite 2T extra tokens — JJitsev · 2026-07-26
- A simple “ask clarifying questions first” prompt gets Claude to surface missing assumptions — Commercial-Most3081 · 2026-07-26
- User says Claude Opus 5 feels more forgiving and “common-sense” than GPT-5.6 Sol — dejavucoder · 2026-07-26
- Frontier models are hallucinating less about biology, but using terms more loosely — owl_posting · 2026-07-26
- This week’s open-weight round: 1M-token Laguna S 2.1, 314B Motif-3-Beta, and more — rasbt · 2026-07-26