Training data is the army: noisy translations make multilingual LLMs lose what makes a language unique
yoavgo · x · 2026-10-06
Presented at COLM 2026, the MameLoshnLM work borrows the Yiddish saying "a language is a dialect with an army" to argue that for LMs, training data is that army: when native texts are overshadowed by noisy translations, multilingual LLMs can look fluent while missing what makes a language unique. Yoav Goldberg amplified the thread.
More from Research
- Proactivity-Gym: 23-agent study shows correct work alone can't sustain user trust — minnesotanlp · 2026-10-06
- SearchJev makes search-agent decisions 5x faster with 41-74% lower calibration error — Congfeng Cao · 2026-10-06
- LLM autoresearch builds AutoSella, cutting DFT force calls to 40-77% of Sella's — Artem Tsypin · 2026-10-06
- ID-Forcing extends short-horizon video diffusion models to minute-scale generation — SeoulNatlUniv · 2026-10-06
- Self-generated feedback destabilizes test-time training, causally decomposed at 128K tokens — KAUST · 2026-10-06
- PFLM: a 300M model pretrained on zero real languages learns languages in context — Lennart Carstens-Behrens · 2026-10-06