Abliterated LLMs Produce 3-4x More False Positives in Bug Hunting

Experiments reproducing FreeBSD kernel CVE detection show abliterated (uncensored) models say 'yes' far more often, inflating false positives 3-4x while missing real bugs—suggesting de-alignment erodes the model's skepticism.

2026-09-03 ~ 2026-09-04 · 2 related posts