Black-Box Attack Bypasses Safety Tuning in T2I Models, Reviving Erased Concepts
pinyuchenTW · x · 2026-08-02
A study accepted at COLM 2026 (ICER) reveals that safety-tuned text-to-image (T2I) models, despite having dangerous concepts like nudity supposedly erased, remain vulnerable to black-box attacks. Attackers can extract these concepts without needing gradients, agents, or auxiliary VLMs, as the models simply remember what worked previously.
More from Safety
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- Debating 'doomsaying for profit' in AI industry — trevposts · 2026-08-24