New Benchmark for Multilingual Moral Reasoning
launch · hf · 2026-07-15
This paper presents three contributions regarding moral reasoning within multilingual and cultural contexts:
- MCLASH: A multilingual moral decision benchmark emphasizing the cultural contexts and social norms embedded in different languages, rather than simply translating English prompts.
- MET: A two-step prompting method. It first selects relevant cultural and situational "grounds" from expert-filtered theoretical foundations in psychology and philosophy, and then reasons in the user's native language.
- MET-D: Incorporates self-distillation training in the second step, requiring no additional human annotation or supervision from stronger models.
Experiments show that MET-D improves performance across multiple models: averaging a 3.71 macro-F1 increase on MCLASH and a 4.23 point increase on MMoralExceptQA. Notably, Qwen3-8B achieved the highest improvement of 12.94 on Malay MCLASH. The authors also found that MET-D boosts native language reasoning by an average of 62.13 points, and that there are systematic differences in the "moral grounds" applicable to different cultures.
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22