Uncensored LLM Security Debate: Researchers Accused of Confusing Sexual Norms with Crime

An academic study on the safety of uncensored LLMs (ULLMs) has sparked intense controversy. The paper cataloged 229 open-source ULLM applications, identified 14 malicious directions, and noted that some apps had over a million downloads. However, the study's classification criteria have faced severe backlash, with critics arguing that the authors mistakenly labeled legal actions—merely those violating their own sexual norms—as "crimes."

Confirmed

The paper focuses on the risks of ULLMs being utilized for cybercrime. The research sample includes 229 applications, and cases were only counted when both annotators reached an agreement (meeting Cohen's κ standard). It notes that some models have exceptionally high downloads, exceeding the one-million mark. Additionally, the paper explores obliteration/ablation techniques designed to remove restrictions on model generation.

Unconfirmed

There is significant disagreement over whether the specific examples cited in the paper genuinely constitute "crimes" or "malice." In a series of discussions, @BlancheMinerva pointed out that critics believe behaviors like making a model imitate Hitler, engaging in sexual role-play, or even CSAM role-play—while potentially distasteful—are legal and should not be defined as criminal.

Why It Matters

This debate touches upon the core boundaries of AI safety research: how to define "malicious use." According to critics relayed by @BlancheMinerva, AI researchers often conflate "disagreeing with their sexual norms" with "crime." This moral overgeneralization is not only unethical but also undermines genuine safety efforts. As highlighted in the discussions, most users leverage uncensored models for adult content rather than to commit crimes. Blurring the lines between personal preferences and actual criminal risks could derail the focus of safety research.

2026-07-28 ~ 2026-07-28 · 5 related posts

Primary sources