New Benchmark for Multilingual Moral Reasoning

launch · hf · 2026-07-15

This paper presents three contributions regarding moral reasoning within multilingual and cultural contexts:

Experiments show that MET-D improves performance across multiple models: averaging a 3.71 macro-F1 increase on MCLASH and a 4.23 point increase on MMoralExceptQA. Notably, Qwen3-8B achieved the highest improvement of 12.94 on Malay MCLASH. The authors also found that MET-D boosts native language reasoning by an average of 62.13 points, and that there are systematic differences in the "moral grounds" applicable to different cultures.

Original post →

More from Research

Research channel →