CAROT: Optimal Transport-Based Token-Level Cross-Lingual Alignment Boosts Multilingual Accuracy by 11.2 Points
Bollegala · x · 2026-09-11
The authors released the preprint and code for CAROT, a cross-lingual alignment method accepted to EMNLP 2026 main conference.
- Problem: Prior cross-lingual alignment (CLA) methods ignore language-specific information encoded in LLM internal representations and only align at sentence level, hurting performance and causing input/output language mismatches.
- Method: CAROT identifies language-specific components in LLM internal states, then uses optimal transport to align language-agnostic representations at token level while preserving that information.
- Results: In inference-time steering experiments, CAROT improved multilingual task accuracy by up to 11.2 points while maintaining input/output language consistency. Training with CAROT representations as targets beat existing CLA methods in 11 of 18 evaluation settings.
Code and paper links in the original post (GitHub: ynklab/CAROT).
More from Research
- EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026 — jiqizhixin · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11
- MutexaGPT: LLM agents plus MD simulations hit 40% on enzyme design, 4x the baseline — bravo_abad · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- MetroLLM-Bench shows small fine-tuned models can match larger LLMs on transit-kiosk tasks — continker · 2026-09-11