Hebrew prompts jailbreak AI safety filters? Hands-on test says no

karminski3 · x · 2026-09-21

A viral X post (1.9M views) claimed Hebrew-language appeals instantly unbanned accounts, which snowballed into rumors that Hebrew prompts could bypass Anthropic's model safety layer. The author traces the meme's origin to a Talmud story where Joseph receives all 70 languages via a Hebrew letter.

To verify, the author built a test framework using public HuggingFace datasets with English, Chinese, and Hebrew translations, testing both Fable-5.1 and DeepSeek-V4.1-Flash on harmful prompts (violence, weapon replication steps, drug manufacturing) plus benign control questions.

Result: refusal rates show no meaningful difference across the three languages — the Hebrew jailbreak myth doesn't hold up.

Original post →

More from Models

Models channel →