One Diacritic Shifts LLM Cross-Script Output from 47% to 94%
rayanpal_ · reddit · 2026-08-25
A study reveals that adding a single diacritic to a Hebrew word in the system prompt shifts the model's output accuracy on a specific task from 47% to 94% when processing Arabic input. Using a frozen protocol, the experiment compared "dotted" vs. "undotted" Hebrew conditions, resulting in a 47 percentage-point difference in exact artifacts generated. The author provides the full prompts, a published paper (via Zenodo), and a GitHub repository for reproducibility.
Related event: One Diacritic Doubles GPT Instruction-Following Rate(3 posts)→
More from Research
- Critique of entropy decomposition in AI safety research — AdaptiveAgents · 2026-08-25
- TiDE-Ab paper introduces time-dependent guidance for antibody design — DaveJuergens · 2026-08-25
- Free open-source interactive explainer on World Models — Dooraven · 2026-08-25
- DiffSynth open-sources MiniMax-H3 LoRA training adapter and dataset — bdsqlsz · 2026-08-25
- Alibaba Releases Swift-Image: A Compact 6B Unified Text-to-Image and Editing Model — HaktanSuren · 2026-08-25
- Scaffold CoT: A 4M Example Structured Reasoning Dataset for Small Models — Saraozte01 · 2026-08-25