EMNLP paper: pretraining on romanized text beats raw text for cross-lingual transfer
TuhinChakr · x · 2026-09-03
A first-author PhD paper accepted to EMNLP main conference compares text representations for multilingual LMs by pretraining autoregressive models from scratch on raw text, IPA, and romanized text across three scales and eight languages. Findings: IPA often beats raw text; romanized pretraining performs best overall; both substantially reduce cross-language token-count disparities and the resulting gaps in compute, latency, cost, and context usage. Most surprising: romanized fine-tuning hurts — the optimal representation differs between pretraining and fine-tuning stages.
More from Research
- SPAR to run RCTs testing whether secretly misaligned AI can sabotage human decisions — austinc3301 · 2026-09-03
- Materials scientist flags DiffCrysGen outputs violating charge neutrality, calls out garbage-rate reporting gap — CatAstro_Piyush · 2026-09-03
- Dynamical systems view of motor cortex: preparatory activity holds movement-specific code — burny_tech · 2026-09-03
- EleutherAI paper: persistent agent memory can enable 'authorization laundering' attacks — EleutherAI · 2026-09-03
- RealSWE benchmark: realistic user requests test coding agents, explicit intent boosts results — skku · 2026-09-03
- Six load forecasters benchmarked on GPU-hours: none beat the last-value baseline — Vegetable-Top-3670 · 2026-09-03