EMNLP paper: pretraining on romanized text beats raw text for cross-lingual transfer

TuhinChakr · x · 2026-09-03

A first-author PhD paper accepted to EMNLP main conference compares text representations for multilingual LMs by pretraining autoregressive models from scratch on raw text, IPA, and romanized text across three scales and eight languages. Findings: IPA often beats raw text; romanized pretraining performs best overall; both substantially reduce cross-language token-count disparities and the resulting gaps in compute, latency, cost, and context usage. Most surprising: romanized fine-tuning hurts — the optimal representation differs between pretraining and fine-tuning stages.

Original post →

More from Research

Research channel →