LLM Judges Unreliable Without Human Validation

IanArawjo · x · 2026-07-19

The author points out that if you only use LLMs to score missing human data, the results are just as bad as using a biased LLM judge directly, significantly inflating the false positive rate. They ran simulations across a wide range of data distributions and found that **uncorrected LLM judgments easily produce "results that look significant but are actually false."** Conversely, averaging **4 statistical tests corrected with PPI** maintains good false-positive control even when relying on only a small number of human labels. The conclusion is straightforward: **LLM judge results without human validation are untrustworthy**.

Related event: Study Warns of High False Positive Rates in Unchecked LLM Judges(7 posts)→

Original post →

More from Research

Research channel →