Xiaohongshu Releases CULTURE-MT: Top LLMs Score Under 40% in Social Media Cultural Translation
小红书技术REDtech · wechat · 2026-08-07
Xiaohongshu, alongside Zhejiang and Fudan Universities, introduced CULTURE-MT, the first benchmark for translating Chinese-English social media notes, aiming to solve the difficulty of conveying cultural intent and emotional resonance in traditional machine translation.
- Evaluation Standard & Tool: The team proposed a new "Cultural Effectiveness" metric and trained an auto-evaluation model JUDGER (based on Qwen3-32B, 86% accuracy) to replace traditional metrics like BLEU that are blind to cultural nuances.
- Key Findings: Testing 15 major models revealed that even the best closed-source models achieve only around 38% in perfect cultural translation. Meanwhile, fine-tuning an 8B model with cultural effectiveness guidance boosted its perfect score rate from 4% to 28%, breaking the scale ceiling.
- Open Source & Hiring: Accepted by ICML 2026, the dataset is open-sourced. The post also includes job postings for multilingual translation and LLM algorithm engineers at Xiaohongshu.
More from Research
- Local MiniMax H3 Video Motion LoRA Training Tested: Runs on 16GB VRAM — ashishsanu · 2026-08-07
- 3D Printed Robot Arm Picks Up Laundry Trained on Single GPU in 3.5 Hours — mishig25 · 2026-08-07
- Study: AI Assistants Are 'Myopically Helpful,' Hindering Independent Thinking — xuanalogue · 2026-08-07
- AI+Physics hybrid framework: Optimizing climate simulation compute with AI predictions — bravo_abad · 2026-08-07
- WeirdML v2 Benchmark Released: Tasks Expanded to 19, Clear Cost-Performance Scaling — yacineMTB · 2026-08-07
- MameLoshnLM: First Open-Source 8B LLM and Benchmark for Yiddish — Yiddish-NLP · 2026-08-07