Study on 229 Uncensored LLMs Sparks Debate on Safety Boundaries
An academic paper cataloging 229 uncensored LLM (ULLM) applications has sparked intense controversy within the AI community. The study identified 14 malicious directions and noted that some apps exceeded a million downloads. However, critics argue that researchers misclassified legal behaviors that merely violate their personal sexual norms as 'crimes,' a moral overgeneralization that could ultimately harm genuine AI safety efforts.
Confirmed
- The paper focuses on the application of uncensored LLMs, with a study sample of 229 apps. Cases were only counted when two annotators agreed (reaching Cohen's κ standard for consistency).
- The paper points out that some model applications have extremely high downloads, exceeding the one million mark.
- The paper discusses obliteration/abliteration techniques. Originally intended to remove content generation restrictions, these models could be repurposed for malicious uses.
Unconfirmed
There is severe disagreement over whether the specific cases cited in the paper actually constitute 'crimes' or 'malicious intent.' As highlighted by @BlancheMinerva in the discussions, critics argue that behaviors like making a model imitate Hitler, or engaging in sexual role-play—even CSAM role-play—while potentially offensive, are legal activities and should not be defined as crimes.
Why it matters
This debate touches the core boundary of AI safety research: how to accurately define 'malicious use.' According to critiques relayed by @BlancheMinerva, AI researchers often equate 'disagreeing with one's sexual norms' with 'crime,' which is both unethical and misleading. @PMinervini emphasized that most people use uncensored models for pornography, not to commit crimes. AI-generated porn is not the most concerning risk; automated cybercrime and fraud pose real safety threats. Blurring the lines between personal preferences and actual criminal risks could derail the focus of safety research.
2026-07-28 ~ 2026-07-28 · 7 related posts
Primary sources
- Study catalogs 229 open-source ULLM apps and flags 14 as malicious — BlancheMinerva ·
- AI safety debate: cybercrime automation, not smut, is the real risk — PMinervini ·
- Researchers clash over whether uncensored models are about porn or cybercrime — BlancheMinerva ·
- [source] Researchers clash over whether uncensored models are about porn or cybercrime — BlancheMinerva · 2026-07-28
- A short reply says most people want porn, not crime — BlancheMinerva · 2026-07-28
- Paper says uncensored LLMs are being used for cybercrime, with some models downloaded millions of times — BlancheMinerva · 2026-07-28
- Post argues model roleplay examples are legal, not criminal — BlancheMinerva · 2026-07-28
- [source] Study catalogs 229 open-source ULLM apps and flags 14 as malicious — BlancheMinerva · 2026-07-28
- [source] AI safety debate: cybercrime automation, not smut, is the real risk — PMinervini · 2026-07-28
- AI safety papers keep conflating sexual content with actual criminal abuse — BlancheMinerva · 2026-07-28