Why Mann-Whitney U Test Falls Short for Likert Data in LLM Evaluations

IanArawjo · x · 2026-08-03

A researcher discovered that when handling 1-5 Likert scale data, the Mann-Whitney U (MWU) test becomes too conservative due to data discreteness (ties) and rounding, sacrificing statistical power.

Currently, Prediction-Powered Inference (PPI) corrected rank-based tests remain unimplemented. Because these tests do not assume normality like t-tests, their gains in saving human labels and boosting power are inherently limited. In contrast, the Wilcoxon signed-rank test does not suffer from this specific MWU flaw.

Original post →

More from Research

Research channel →