LLM Judges Unreliable Without Human Validation

IanArawjo · x · 2026-07-19

The author points out that if you only use LLMs to score missing human data, the results are just as bad as using a biased LLM judge directly, significantly inflating the false positive rate.

They ran simulations across a wide range of data distributions and found that uncorrected LLM judgments easily produce "results that look significant but are actually false." Conversely, averaging 4 statistical tests corrected with PPI maintains good false-positive control even when relying on only a small number of human labels. The conclusion is straightforward: LLM judge results without human validation are untrustworthy.

Related event: Study Warns: Unchecked LLM Judges Yield High False Positives(7 posts)→

Original post →

More from Research

Research channel →