Xiaomi's Open Multilingual Translation Model Outperforms Proprietary Baselines
xiaomi-research · hf · 2026-08-12
Xiaomi's research team proposes a reference-free post-training method for open-source LLMs to enhance multilingual machine translation.
The approach utilizes Group Relative Policy Optimization (GRPO) with reference-free quality rewards alongside checkpoint interpolation. Experiments demonstrate that open-source models optimized with this post-training method not only surpass other strong open-source models but also defeat mainstream proprietary baselines.
More from Models
- FLUX 3 Video hits #2 on Arena, just 16 points behind Gemini — arena · 2026-08-12
- Testing MiniMax R2V b20-49 Hybrid: Cleaner Audio and Better Voice Cloning — Foreforks · 2026-08-12
- Anthropic to Embed Invisible Watermarks in Claude Text Globally Under EU AI Act — TinfoilTricorn · 2026-08-12
- AI Safety Guardrails Block Proof of Irreducibility for Stern Polynomials — chaumian · 2026-08-12
- DeepSeek Prefix Cache Hacks: Cut Agent Token Costs by 90% to $0.005/Task — BodybuilderLost328 · 2026-08-12
- Local MoE Benchmark: NVIDIA Lightning Outruns Qwen by 2.5x — parepeg · 2026-08-12