Microsoft Research: Action Policies Outperform Final Answers in Multilingual Agent Evaluation

dair_ai · x · 2026-08-13

A new study from Microsoft Research argues that traditional evaluation methods for multilingual agents, which only compare final answers, are limited. Instead, the model's action trajectory is the critical object for measuring cross-lingual consistency.

The team conducted 2.38 million rollouts across 6 benchmarks and 41 languages. They identified five confounding factors in raw trace similarity: short traces scoring higher, empty traces scoring perfectly, chance agreement among unrelated traces, reproducibility capping the gap, and intra-language variability.

After controlling for these confounds, the study found that frontier models maintain 71% to 73% of their action policy consistency across different languages.

Original post →

More from coding & agent

coding & agent channel →