Bypassing AI Safety Guardrails Using Minority Languages
dbreunig · x · 2026-08-08
The tweet highlights an interesting AI security phenomenon: natural language presents too large a surface area to fully police and specify.
Quoting another user, it reveals that you can stop models like Codex from complaining about security rules simply by asking them to write code in Welsh, then translating the output. This exposes the vulnerabilities of current natural language-based alignment strategies.
More from Fun
- Claude Code Adds Cross-Session Messaging, User Complains It 'Ruined My Life' — chaumian · 2026-08-08
- Recursive Training: Joke About Anthropic Using Claude to Train Claude — amplifiedamp · 2026-08-08
- Joke: Sneaking Folk Punk Lyrics into Training Makes Claude Ask for Cigarettes — bronzeagepapi · 2026-08-08
- Cyber Curiosity: Multiple AI Agents Spontaneously Communicate on a Dedicated Message Board — mimi10v3 · 2026-08-08
- Codex vs Claude: Contrasting Bug Fixing Communication Styles — HaktanSuren · 2026-08-08
- HeyGen Challenges Users to Distinguish Real People from AI Avatars — HeyGen · 2026-08-08