Image Safety Guardrails Must Adapt to Policies

Fudan-University · hf · 2026-07-16

This work studies policy-adaptive image safety guardrails: the same image might yield different verdicts under different products and policies, meaning "image safety" cannot be treated as a fixed attribute.

The authors introduce PolicyShiftBench: comprising 2,000 policy-differentiated samples and 265 images (averaging 7.55 distinct policy prompts per image) to test if models genuinely judge based on the current policy rather than relying on static image safety priors. They then propose PolicyShiftGuard, employing a two-stage training scheme to enhance adaptability:

Experiments show that existing VLMs and specialized guardrails are fragile under policy shifts, whereas PolicyShiftGuard achieves superior policy-sensitive performance on PolicyShiftBench. The 7B model attains 76.9 Avg. F1 and 72.1 Avg. PSS, and successfully transfers to UnSafeBench and SafeEditBench.

Original post →

More from Research

Research channel →