SafeAtlas-VL: Graded Multimodal Safety Dataset and Guard Models Hit SOTA
SJTU · hf · 2026-09-01
SJTU's SafeAtlas-VL presents a large-scale multimodal safety dataset with graded risk labels instead of binary flags, and trains guard models that predict continuous safety scores. The approach achieves state-of-the-art generalization, moving multimodal safety beyond binary safe/unsafe classification.
More from Safety
- Can Interpretationism Explain Beliefs and Deception in AI Agents? — raphaelmilliere · 2026-09-01
- CNRS researchers forced to use Mistral, banned from OpenAI/Anthropic models — eliebakouch · 2026-09-01
- Prediction: 99% of Researchers Will Work on Safety-Related Roles — maksym_andr · 2026-09-01
- Volcengine releases AgentSentry for unified enterprise Agent security management — 火山引擎 · 2026-09-01
- HuggingFace incident reveals covert channels need only simple HTTP ambiguity — orionintx · 2026-09-01
- Public Safety AI: How Peregrine Uses Agents to Solve Cold Cases — Training Data (Sequoia) · 2026-09-01