Hebrew prompts jailbreak AI safety filters? Hands-on test says no
karminski3 · x · 2026-09-21
A viral X post (1.9M views) claimed Hebrew-language appeals instantly unbanned accounts, which snowballed into rumors that Hebrew prompts could bypass Anthropic's model safety layer. The author traces the meme's origin to a Talmud story where Joseph receives all 70 languages via a Hebrew letter.
To verify, the author built a test framework using public HuggingFace datasets with English, Chinese, and Hebrew translations, testing both Fable-5.1 and DeepSeek-V4.1-Flash on harmful prompts (violence, weapon replication steps, drug manufacturing) plus benign control questions.
Result: refusal rates show no meaningful difference across the three languages — the Hebrew jailbreak myth doesn't hold up.
More from Models
- Chinese open-source labs explode on OpenRouter: Moonshot +2425%, Z.ai +1925%, DeepSeek +1000% — FinanceYF5 · 2026-09-21
- Yacine rants: 'How could they ship an LLM that is so dog shit at programming?' — yacineMTB · 2026-09-21
- Yacine: coding LLMs produce 'total complex garbage' — I still read every line — yacineMTB · 2026-09-21
- Researcher: LLMs write convincing related work, but convincing isn't comprehensive — lucacarlone1 · 2026-09-21
- OpenAI's secret technique for upcoming Astra model sparks security concerns — keviv9 · 2026-09-21
- Jev + graphical models: zero-shot probabilistic factors could reshape reasoning systems — fdellaert · 2026-09-21