HarmProfile Benchmark: Harmfulness and Diversity Rise with Model Capability
Zhouyuan Ma · hf · 2026-08-19
HarmProfile is a benchmark dataset characterizing frontier LLM safety failures through content analysis. It reveals that harmfulness and diversity increase with model capability.
More from Safety
- xAI Updates Grok 4.6 Model Card, Revises Evaluations — Miles_Brundage · 2026-08-19
- Berkeley Professor's Pro-SAT Op-Ed Flagged as 33% AI-Generated — lpachter · 2026-08-19
- NHS Pauses Access to National-Scale 60M-Patient AI Model — zakkohane · 2026-08-19
- Anthropic's AI watermarking shows disregard for writing quality — nikvassev · 2026-08-19
- Critical vulnerabilities found in conference review system HotCRP — moyix · 2026-08-19
- Fathom CEO: Companies grading own homework risks unsafe AI adoption — ghadfield · 2026-08-19