Open-source tool 'heretic' auto-strips model refusals: Gemma 3 12B drops from 97 to 3 refusals per 100 prompts

thisguyknowsai · x · 2026-09-05

A free open-source tool called heretic removes censorship layers from local open-weight models with a single command — no jailbreak prompts needed. It edits the model directly: it searches for modifications that reduce refusals while preserving other abilities, tests changes, measures refusal rates, checks drift in normal responses, and lets you save and chat with the modified model. In the developer's test on Google's Gemma 3 12B, refusals fell from 97 out of 100 harmful test prompts to just 3. Caveats: this is one reported benchmark, and fewer refusals don't mean better or more accurate answers.

Original post →

More from Models

Models channel →