Why Mann-Whitney U Test Falls Short for Likert Data in LLM Evaluations
IanArawjo · x · 2026-08-03
A researcher discovered that when handling 1-5 Likert scale data, the Mann-Whitney U (MWU) test becomes too conservative due to data discreteness (ties) and rounding, sacrificing statistical power.
Currently, Prediction-Powered Inference (PPI) corrected rank-based tests remain unimplemented. Because these tests do not assume normality like t-tests, their gains in saving human labels and boosting power are inherently limited. In contrast, the Wilcoxon signed-rank test does not suffer from this specific MWU flaw.
More from Research
- DMC-Optim: A New Benchmark for Training AI to Write Faster Code — burkov · 2026-08-03
- ExtractBench: First Comprehensive Benchmark for Enterprise Document Extraction — Boyang Zhang · 2026-08-03
- Meshy T2: Fast Native Mesh Generation Framework via Flow Matching — Jiale Xu · 2026-08-03
- EVR Reward Model: Breaking the Consistency Bottleneck in Multi-Reference Image Editing — Yingmao Miao · 2026-08-03
- QQWorld: Fixing Heavy-Tailed Deviations in World Models with Quantile-Quantile Matching — Zhoushun Yu · 2026-08-03
- N_0-VTLA: First VTLA Foundation Model Pretrained on Tactile Data at Scale — NeoteAIEmbodied · 2026-08-03