Uncensoring a Model Also Removes Its Skepticism: 3-4x False Positives

dyn___ · x · 2026-09-03

A 17-minute technical blog post reports an empirical finding: while testing local open-weight models against a known FreeBSD kernel CVE, the author noticed abliterated (uncensored) builds say "yes" far more often.

Takeaway: "Remove the refusal, and the skepticism goes with it" — abliteration damages critical judgment, not just refusals.

Related event: Abliterated LLMs Produce 3-4x More False Positives in Bug Hunting(2 posts)→

Original post →

More from Research

Research channel →