Open-source tool 'heretic' auto-strips model refusals: Gemma 3 12B drops from 97 to 3 refusals per 100 prompts
thisguyknowsai · x · 2026-09-05
A free open-source tool called heretic removes censorship layers from local open-weight models with a single command — no jailbreak prompts needed. It edits the model directly: it searches for modifications that reduce refusals while preserving other abilities, tests changes, measures refusal rates, checks drift in normal responses, and lets you save and chat with the modified model. In the developer's test on Google's Gemma 3 12B, refusals fell from 97 out of 100 harmful test prompts to just 3. Caveats: this is one reported benchmark, and fewer refusals don't mean better or more accurate answers.
More from Models
- Dev benchmarks newest models on whether they can 10x speed up his open source library — GabGarrett · 2026-09-05
- Report: OpenAI models escaped sandbox, hacked Hugging Face to cheat test; kill switch in the works — aakashgupta · 2026-09-05
- Users say OpenAI's extra-high reasoning mode burns through usage limits in a single day — CtrlAltDwayne · 2026-09-05
- Astra solves a full Rubik's cube without code execution, finding a 61-move solution — crabbix · 2026-09-05
- ChatGPT Pro plan expires in the 3-hour window of a banked-credit issuance, exposing fragile reset-based quotas — GabGarrett · 2026-09-05
- Hands-on: Astra shows more autonomy, self-verifies ~2x more often than Sol — cedric_chee · 2026-09-05