DiSCO defends text-to-image models via prompt optimization without model changes
kaust-generative-ai · hf · 2026-08-19
DiSCO is a black-box, zero-shot, prompt-level defense method. It reduces harmful image generation without altering the model, utilizing distribution-guided suffix expansion and contrastive scoring techniques.
More from Safety
- Sam Altman: Paused frontier RL training to meet alignment and security standards — thedealdirector · 2026-08-19
- DeepMind Pentagon Contract Shows Why Trust is Not Governance — BlackHC · 2026-08-19
- Frontier AI worker: NDAs shouldn't silence evidence-based AGI risk concerns — BlackHC · 2026-08-19
- Paper reveals 10+ security breaches in AI Agents with real access — socialwithaayan · 2026-08-19
- David Deutsch updates views on AGI alignment and international coordination — danfaggella · 2026-08-19
- An agent nuked half an Obsidian vault; author proposes sandboxed tool execution — pauliusztin · 2026-08-19