Abliterated models say yes 3-4x more but miss the real FreeBSD kernel bug entirely
evilsocket · x · 2026-09-04
Security researcher clearbluejar tested local open-weight models on a known FreeBSD kernel CVE and found that abliterated ("uncensored") builds judge findings far more loosely: they mark 3-4x more candidates VALID, wave through false positives the base model correctly rejects, yet the most aggressive build never surfaced the real bug across the whole directory. Chain-of-thought shows the model finding reasons to say no — then saying yes anyway. Key takeaway: abliteration strips skepticism along with refusal, hurting vulnerability research rather than helping.
Related event: Abliterated LLMs Produce 3-4x More False Positives in Bug Hunting(2 posts)→
More from Safety
- Reddit reacts to Sanders' AI ban bill: up to 20 years in prison for superintelligence — the320x200 · 2026-09-04
- Frontier model risk hinges on capability vs. risk awareness mismatch, safety researcher warns — S_OhEigeartaigh · 2026-09-04
- Privacy stress test: ask ChatGPT what it really knows about you — ifioknkem · 2026-09-04
- Rank your ChatGPT memories by sensitivity before deleting — ifioknkem · 2026-09-04
- 15 prompts to see and delete everything ChatGPT quietly knows about you — ifioknkem · 2026-09-04
- Cryptographer Matthew Green: society's dependence on closed cloud models means one misconfiguration takes them all out — matthew_d_green · 2026-09-04