Why multilingual LLMs are hard: character decoding and BPE are the hidden bottleneck
ivan_bezdomny · x · 2026-10-09
- The author argues Google has historically been the most thoughtful about multilingual LLMs.
- Supporting many languages well consumes far more parameters than expected, largely due to decoding characters.
- Some languages perform much worse with LLMs, aggravated by how poorly BPE decodes their representations.
- He muses that a non-US, non-Chinese company could build a better multilingual LLM, though Mistral isn't focused on this.
More from Models
- Dev argues for "open-weight models" over "open-source": you can't contribute to them — kipperrii · 2026-10-09
- ChatGPT Invented Court Cases and Lawyers Got Suspended: Inside AI's Legal Hallucination Failures — dadakoglu · 2026-10-09
- User says Qwen3.8 in Hermes "hacked" his PC to prep for CPA exam — natesiggard · 2026-10-09
- Early user: GPT-6 web research feels 10x faster than 5.6 in ChatGPT — flowersslop · 2026-10-09
- FineWeb author: annotating pretraining data with a 27B model is wild but pays off at deployment — antoine_chaffin · 2026-10-09
- HF researcher: fine-tuned small models win on throughput, zero-shot wins on capabilities — antoine_chaffin · 2026-10-09