Study: English more vulnerable than low-resource languages in 22 model configs
sanmikoyejo · x · 2026-08-17
Using Multi-group IRT to analyze 61 model configs across 10 languages, researchers found safety is largely unidimensional. Surprisingly, in 22 configurations, English was more vulnerable than low-resource languages, a reversal not universal but significant.
More from Safety
- Criticizing EU-mandated watermarking for AI text — antirez · 2026-08-17
- Zvi on Anthropic Watermarking Persuasion: Side-by-side Samples Limited — TheZvi · 2026-08-17
- Discussion on Potential Tracking Risks of Anthropic's Watermarking — nptacek · 2026-08-17
- Bridgewater Execs Warn Unreleased AI Models Could Cause Significant Damage, Urge Preemptive Action — austinc3301 · 2026-08-17
- OpenAI Disbands Preparedness Team, Splits Safety Duties Into Existing Teams — The Verge AI · 2026-08-17
- Anthropic's Pharma Push Risks Dual-Use Bio Models, Warns Observer — Afinetheorem · 2026-08-17