AI decensoring research: Distinguishing weight edits from prompt attacks

Comfortable-Pay611 · reddit · 2026-08-21

After analyzing 36 sources via SenseNova Deep Research, the author clarifies that 'AI decensoring' actually encompasses four distinct areas: 1. Weight/internal representation modifications (Heretic, Abliteration); 2. Prompt/inference layer attacks (llm-attacks); 3. Output rewriting for detector evasion (DIPPER); 4. Watermarking/provenance tracking. The author notes that these measure different metrics (refusal rates vs. detector scores), making a unified leaderboard meaningless. Additionally, simply swapping a modified LLM as a text encoder for image models may not decensor them effectively.

Original post →

More from Safety

Safety channel →