Bypassing AI Safety Guardrails Using Minority Languages

dbreunig · x · 2026-08-08

The tweet highlights an interesting AI security phenomenon: natural language presents too large a surface area to fully police and specify.

Quoting another user, it reveals that you can stop models like Codex from complaining about security rules simply by asking them to write code in Welsh, then translating the output. This exposes the vulnerabilities of current natural language-based alignment strategies.

Original post →

More from Fun

Fun channel →