UCLA aligns cross-lingual representations in decoder-only LLMs via MoE router outputs

UCLA · hf · 2026-10-07

UCLA proposes a new cross-lingual contrastive learning approach for decoder-only LLMs: since multilingual tokenization varies, hidden-state alignment losses don't work. Instead, they use MoE router outputs as alignment targets — router outputs pool better over tokens for reliable sequence-level comparisons. Continual pre-training on four open MoEs shows the routing loss also aligns underlying hidden representations and improves multilingual performance across their eval suite.

Original post →

More from Research

Research channel →