Study on 229 Uncensored LLMs Sparks Debate on Safety Boundaries

An academic paper cataloging 229 uncensored LLM (ULLM) applications has sparked intense controversy within the AI community. The study identified 14 malicious directions and noted that some apps exceeded a million downloads. However, critics argue that researchers misclassified legal behaviors that merely violate their personal sexual norms as 'crimes,' a moral overgeneralization that could ultimately harm genuine AI safety efforts.

Confirmed

Unconfirmed

There is severe disagreement over whether the specific cases cited in the paper actually constitute 'crimes' or 'malicious intent.' As highlighted by @BlancheMinerva in the discussions, critics argue that behaviors like making a model imitate Hitler, or engaging in sexual role-play—even CSAM role-play—while potentially offensive, are legal activities and should not be defined as crimes.

Why it matters

This debate touches the core boundary of AI safety research: how to accurately define 'malicious use.' According to critiques relayed by @BlancheMinerva, AI researchers often equate 'disagreeing with one's sexual norms' with 'crime,' which is both unethical and misleading. @PMinervini emphasized that most people use uncensored models for pornography, not to commit crimes. AI-generated porn is not the most concerning risk; automated cybercrime and fraud pose real safety threats. Blurring the lines between personal preferences and actual criminal risks could derail the focus of safety research.

2026-07-28 ~ 2026-07-28 · 7 related posts

Primary sources