Paper: Why Safety Guardrails Degrade Across Languages
sanmikoyejo · x · 2026-08-17
This paper investigates why LLM safety degrades in non-English languages. It introduces a Multi-Group IRT framework to decompose aggregate scores into language-agnostic robustness, intrinsic hardness, processing difficulty, and cross-lingual safety gaps.
Key Findings:
- Unidimensionality: Safety mechanisms are largely shared across harm types.
- Reversal: English is more vulnerable than low-resource languages in 22 of 61 configurations.
- Diagnosis over Ranking: Decomposition helps move from leaderboards to actionable diagnostics.
- High-risk Prompts: High-gap prompts cluster in physical harm categories (theft, weapons) and lower-resource languages.
More from Safety
- Criticizing EU-mandated watermarking for AI text — antirez · 2026-08-17
- Zvi on Anthropic Watermarking Persuasion: Side-by-side Samples Limited — TheZvi · 2026-08-17
- Discussion on Potential Tracking Risks of Anthropic's Watermarking — nptacek · 2026-08-17
- Bridgewater Execs Warn Unreleased AI Models Could Cause Significant Damage, Urge Preemptive Action — austinc3301 · 2026-08-17
- OpenAI Disbands Preparedness Team, Splits Safety Duties Into Existing Teams — The Verge AI · 2026-08-17
- Anthropic's Pharma Push Risks Dual-Use Bio Models, Warns Observer — Afinetheorem · 2026-08-17