Image Safety Guardrails Must Adapt to Policies
Fudan-University · hf · 2026-07-16
This work studies policy-adaptive image safety guardrails: the same image might yield different verdicts under different products and policies, meaning "image safety" cannot be treated as a fixed attribute.
The authors introduce PolicyShiftBench: comprising 2,000 policy-differentiated samples and 265 images (averaging 7.55 distinct policy prompts per image) to test if models genuinely judge based on the current policy rather than relying on static image safety priors. They then propose PolicyShiftGuard, employing a two-stage training scheme to enhance adaptability:
- RP-SFT: Randomized Policy Supervised Fine-Tuning
- BP-Adapt: Contrastive learning using paired "pass/block" boundary samples
Experiments show that existing VLMs and specialized guardrails are fragile under policy shifts, whereas PolicyShiftGuard achieves superior policy-sensitive performance on PolicyShiftBench. The 7B model attains 76.9 Avg. F1 and 72.1 Avg. PSS, and successfully transfers to UnSafeBench and SafeEditBench.
More from Research
- SUFLECA shows NOC-based correspondence can improve CAD-to-image alignment — ducha_aiki · 2026-07-21
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21