UCLA aligns cross-lingual representations in decoder-only LLMs via MoE router outputs
UCLA · hf · 2026-10-07
UCLA proposes a new cross-lingual contrastive learning approach for decoder-only LLMs: since multilingual tokenization varies, hidden-state alignment losses don't work. Instead, they use MoE router outputs as alignment targets — router outputs pool better over tokens for reliable sequence-level comparisons. Continual pre-training on four open MoEs shows the routing loss also aligns underlying hidden representations and improves multilingual performance across their eval suite.
More from Research
- Datology AI open-sources Zephon, a deterministic on-the-fly data loader built for massive ablations — pratyushmaini · 2026-10-07
- MIT leaves a coding agent alone on a digital island for 30 hours to study if machines play — phillip_isola · 2026-10-07
- GPT-7 'proves' multiplication is 0.0000000000163% faster, mathematicians react — jxmnop · 2026-10-07
- Symmetrix-XL: open-source engine simulates 10M atoms on a single GPU — CatAstro_Piyush · 2026-10-07
- Swapping just the system prompt measurably changes harness benchmark results — yb2698 · 2026-10-07
- Best-scoring harness combo burns nearly 1.5x more tokens than Qwen Code — yb2698 · 2026-10-07