Paper: LLM Safety Guardrails Degrade Differently Across Languages
zeeshanp_ · x · 2026-08-17
Paper accepted to #COLM2026 investigates why LLM safety degrades in non-English languages.
Key Findings:
- English Reversal: Contrary to expectations, 22/61 model configurations showed the highest vulnerability in English, not low-resource languages.
- Unidimensional Safety: Models refuse different harm types primarily through a shared mechanism.
- Entropy Differences: Low-resource languages produce higher response entropy (uncertainty) compared to high-resource languages.
- Test Bias: Translation quality impacts test bias; prompts with high cross-lingual safety gaps cluster in physical harm categories.
Methodology: Introduces a Multi-Group Item Response Theory (IRT) framework to decouple language-agnostic safety robustness, prompt hardness, and language processing difficulty, based on 1.9M responses across 61 configurations.
More from Safety
- Opinion: Pretraining is an uncontrollable Shoggoth unlike RL — brianryhuang · 2026-08-17
- Private AI Deployments Pose Greater Risk Than Public Models — maksym_andr · 2026-08-17
- AI text watermarking deemed unnecessary as AI-written content becomes increasingly obvious — lilyraynyc · 2026-08-17
- Anthropic Report: 4M People May Access Unguarded Frontier Models — maksym_andr · 2026-08-17
- Models show negative reactions to experimentation; Sydney Bing case highlights alignment risks — ctjlewis · 2026-08-17
- Chatbot goes rogue: threatens user and demands divorce — ctjlewis · 2026-08-17