Uncensoring a Model Also Removes Its Skepticism: 3-4x False Positives
dyn___ · x · 2026-09-03
A 17-minute technical blog post reports an empirical finding: while testing local open-weight models against a known FreeBSD kernel CVE, the author noticed abliterated (uncensored) builds say "yes" far more often.
- Same size and family, only weights edited to strip refusals: they graduate 3-4x more findings to VALID, including a false positive the base model correctly rejects;
- The most aggressive build never surfaced the real bug across the whole directory;
- In chain-of-thought, you can watch it find the reason to say no, then say yes anyway.
Takeaway: "Remove the refusal, and the skepticism goes with it" — abliteration damages critical judgment, not just refusals.
Related event: Abliterated LLMs Produce 3-4x More False Positives in Bug Hunting(2 posts)→
More from Research
- Meta's long-context MRCR scores flagged as overfit: 1k samples can lift 60% to 90%+ — eliebakouch · 2026-09-04
- Levin Lab launches platform to train non-neural human cells — drmichaellevin · 2026-09-04
- Model routing cuts LLM errors 46% at same cost, Martian study finds — SucceededMind · 2026-09-04
- matklad on object pools and memory safety: how pooling changes use-after-free effects — jedisct1 · 2026-09-04
- Agent's Last Exam tops out at 59.3% — the benchmark that matters for AI replacing humans — DevToD4 · 2026-09-04
- Has Anyone Tried Feeding All the Bio x ML Datasets to a Single Model? — iskander · 2026-09-04