Black-Box Attack Bypasses Safety Tuning in T2I Models, Reviving Erased Concepts

pinyuchenTW · x · 2026-08-02

A study accepted at COLM 2026 (ICER) reveals that safety-tuned text-to-image (T2I) models, despite having dangerous concepts like nudity supposedly erased, remain vulnerable to black-box attacks. Attackers can extract these concepts without needing gradients, agents, or auxiliary VLMs, as the models simply remember what worked previously.

Original post →

More from Safety

Safety channel →